“Can a 7B model fit on my GPU?” sounds like a parameter-count question. It is actually a weight format, architecture, context, concurrency, and runtime question.

The quick weight floors

7B: 3.26 GiB INT4 · 6.52 GiB INT8 · 13.04 GiB FP16.

13B: 6.05 GiB INT4 · 12.11 GiB INT8 · 24.21 GiB FP16.

Those numbers exclude KV cache, activations, kernels, workspace, and runtime headroom.

The 32K cache can outweigh the quantization win

In the guide’s illustrative GQA profiles, the 7B model uses a 4 GiB 32K BF16 cache and the 13B model uses 5 GiB. With INT4 weights and 20% headroom, the planning targets are 8.71 GiB and 13.26 GiB.

Change only the KV-head count to an MHA profile and the targets jump to 23.11 GiB and 37.26 GiB. The 13B cache alone rises from 5 GiB to 25 GiB.

Open the exact model plan

Replace parameter count, layer count, KV heads, head dimension, context, concurrency, precision, headroom, and usable VRAM in the calculator.

Already subscribed? Put your referral link to work.

Share this GPU worksheet with one teammate. Three confirmed referrals unlock the AI Launch Evidence Checklist automatically.

Not subscribed yet? Join ResearchAudio free →

Keep Reading