
Q4 7B-8B models fit an 8 GB card at moderate context: their files are 4.4 to 4.9 GB. The catch is the KV cache. For Llama 3.1 8B it adds about 1 GB at 8K and about 4 GB at 32K in FP16, which pushes the total past 8 GB. KV cache quantization (q8_0, about half the memory), flash attention, and layer offload keep it usable.
An 8 GB graphics card is one of the most common cards in a gaming PC, and it runs local models well. The question people usually ask is whether a 7B or 8B model fits. It does. The catch is that the model file is only half of the bill.
I covered memory types and hardware tiers in a broader VRAM and RAM guide. This post is narrower: what actually fits in 8 GB of VRAM, and what to change when the total runs past the limit.
The model file is only half your VRAM bill
The short version: A 7B to 8B model at Q4_K_M is a 4.4 to 4.9 GB file, so the weights fit 8 GB with room left over. Everything after that is a decision about context.
Start with file sizes, not parameter counts. The llama.cpp quantize docs give the ratio for Llama 3.1 8B: Q3_K_M at 4.00 bits per weight, Q4_K_M at 4.89, Q5_K_M at 5.70, Q6_K at 6.56, and Q8_0 at 8.50. The same table puts the Q4_K_M file at 4.9 GB on disk, down from 32.1 GB unquantized.
The method checks out against a real download: 8 billion parameters at 4.8944 bits per weight gives 8 x 4.8944 / 8 = 4.89 GB estimated, and the actual Llama 3.1 8B Q4_K_M file is 4.92 GB.
The Q4_K_M sizes that matter on an 8 GB card, from the model cards:
| Model | Q4_K_M file | Fits 8 GB? |
|---|---|---|
| Llama 3.2 3B Instruct | 2.02 GB | yes, easily |
| Mistral 7B Instruct v0.3 | 4.37 GB | yes |
| Qwen2.5 7B Instruct | 4.68 GB | yes |
| Llama 3.1 8B Instruct | 4.92 GB | yes |
| Gemma 2 9B Instruct | 5.76 GB | borderline |
| Qwen2.5 14B Instruct | 8.99 GB | no, over the card itself |
Each card lists the rest of the ladder too. On Llama 3.1 8B it runs 4.92 GB at Q4_K_M, 5.73 GB at Q5_K_M, 6.60 GB at Q6_K, and 8.54 GB at Q8_0, the last of which is bigger than the card before any context loads. Qwen2.5 7B is the same story: 4.68 GB at Q4_K_M, 8.10 GB at Q8_0. And Mistral 7B at 4.37 GB (about 4.83 bits per weight) sits in the same band, with a Q3_K_M at 3.52 GB.
Context is not free: the KV cache grows with every token
The short version: Every token in the conversation adds to a KV cache that lives in the same memory as the weights. For Llama 3.1 8B in FP16 that is about 1 GB at 8K tokens and about 4 GB at 32K, which is what pushes a fitting model over the line.
The Hugging Face Llama 3.1 blog puts it directly: “the cache uses as much memory as the weights when approaching the context length maximum”. Their 8B figures are about 4 GB at INT4 just to load the model, then 0.125 GB of FP16 cache at 1k tokens, 1.95 GB at 16k, and 15.62 GB at 128k.
The cache has a formula, from KV cache basics:
cache bytes = 2 x layers x KV heads x head_dim x bytes per value x tokens
The 2 covers the key and value tensors in each layer. The Transformers docs note that this cache can bottleneck long context, and that Mistral and Gemma 2 use sliding-window attention, where it stops growing at the window.
Llama 3.1 8B has 32 layers, 8 KV heads, and a head dimension of 128, per its config:
- per token: 2 x 32 x 8 x 128 x 2 = 131,072 bytes, or 128 KiB
- 8K tokens: about 1.00 GiB. Estimated total: 4.92 + 1.00 = 5.92 GB
- 32K tokens: about 4.00 GiB. Estimated total: 4.92 + 4.00 = 8.92 GB, past the card
Qwen2.5 7B has 28 layers, 4 KV heads, and the same 128 head dimension, per its config:
- per token: 2 x 28 x 4 x 128 x 2 = 57,344 bytes, or 56 KiB
- 8K tokens: 448 MiB. Estimated total: 4.68 + 0.44 = 5.12 GB
- 32K tokens: about 1.75 GiB. Estimated total: 4.68 + 1.75 = 6.43 GB
The gap is grouped-query attention. Llama 3.1 8B has 8 KV heads against 4 on Qwen2.5 7B, so the Qwen cache is about half the size per token. Fewer KV heads means a longer conversation fits the same memory.
Every total above is an estimate, and none includes compute buffers or desktop usage.
The 8 GB shortlist: what fits at Q4 and what does not
The short version: 3B and 7B to 8B models at Q4_K_M fit at moderate context with room to spare. Gemma 2 9B is borderline. 12B to 14B models, and anything at Q8_0, do not fit without offload.
3B and 7B to 8B at Q4: a comfortable fit
Llama 3.2 3B at 2.02 GB leaves a large context budget, which makes it the safe pick when the card also drives your display. Mistral 7B at 4.37 GB, Qwen2.5 7B at 4.68 GB, and Llama 3.1 8B at 4.92 GB all land at roughly 5.1 to 5.9 GB estimated total at 8K context, leaving a gigabyte or two for everything else. For how small models behave in an actual conversation, see Best Small LLMs for Roleplay on 8 GB of RAM.
Gemma 2 9B at Q4: borderline
The Gemma 2 9B Q4_K_M file is 5.76 GB, leaving roughly 2.2 GB for the cache, compute buffers, and the desktop. At a modest context with a quantized cache it can work. At 32K it will not.
12B to 14B at Q4, or 7B and 8B at Q8_0: no
Qwen2.5 14B at Q4_K_M is 8.99 GB, over 8 GB before a single token of context. Q8_0 is just as far out of reach: 8.10 GB for Qwen2.5 7B and 8.54 GB for Llama 3.1 8B. These can run with layer offload, below, but they do not fit on the card alone.
A short note on Macs
Apple silicon draws no line between memory pools. Metal’s shared storage mode, the default on Apple silicon, is defined as “the CPU and GPU share access to the resource, allocated in system memory”, and MLX adds that “Apple silicon has a unified memory architecture. The CPU and GPU have direct access to the same memory pool.” On a Mac your limit is total unified memory. On a Windows card, 8 GB is 8 GB.
How to run a model that does not quite fit
The short version: Offload some layers to the CPU, quantize the KV cache to q8_0, turn on flash attention, and keep context and batch sizes modest. It runs slower, but it runs.
Partial offload is normal, not a failure. The Ollama FAQ notes that a model goes to the GPU when it fits, that flash attention reduces memory as context grows, and that KV quantization uses about half the memory of f16 at q8_0 and about a quarter at q4_0. It also shows a partial load: ollama ps reporting a split such as 48% on the CPU and 52% on the GPU.
In llama.cpp itself, the server docs list the flags that matter: -ngl or --gpu-layers (how many layers stay in VRAM, a number, auto, or all), -c or --ctx-size (context), -b (default 2048) and -ub (default 512) for batch sizes, and -ctk and -ctv for the KV cache type, default f16 with q8_0 and q4_0 allowed.
With a q8_0 cache, Llama 3.1 8B needs roughly 0.50 GiB at 8K (estimated total about 5.42 GB) and roughly 2.00 GiB at 32K (estimated total about 6.92 GB, tight but under the line). Same model, same card, four times the context, because the cache was cut in half.
The trade is simple. A smaller model fully on the GPU is fast and predictable. A larger model split between GPU and CPU is slower, and how much slower depends on your CPU. For a companion you talk to every day, I would rather run a 7B fully on the card than a 14B that stutters.
The Windows desktop toll: measure it, because nobody publishes it
The short version: There is no official figure for how much VRAM Windows and your display take. Measure idle usage in Task Manager, leave real headroom, and plan around what you find.
I could not find an official Microsoft, Nvidia, or AMD number for the VRAM that Windows and desktop composition cost on a card. The nearest official thing is Ollama’s OLLAMA_GPU_OVERHEAD setting, “Reserve a portion of VRAM per GPU (bytes)”. That setting exists because the overhead is real, but nobody puts a number on it.
So measure it yourself. Close the heavy apps, open Task Manager, and note how much VRAM is in use before any model loads. That number is yours to subtract from 8 GB, and it grows with a browser, a game launcher, or extra monitors. Once a model loads, ollama ps shows whether it is fully on the GPU or split with the CPU.
That is the whole method: sizes from the cards, cache from the formula, then measure your own desktop instead of trusting a guess. If the total is over, the fix is usually a q8_0 cache or a shorter context, and only then a smaller quant. Your machine gets the final vote.
Local Waifu runs these models offline on Windows and Mac, picks a model that fits your card, and keeps the conversation on your own machine. No command line, no flags to get wrong. If the 8 GB card in this post is yours, download it here, let it set up the model, and see what it chose. The free seven-day trial is enough time to see what your hardware can actually hold.
Questions people ask
Does a 7B or 8B model at Q4 fit in 8 GB of VRAM?
Yes at moderate context. Llama 3.1 8B Q4_K_M is a 4.92 GB file plus about 1 GB of cache at 8K context. At 32K the FP16 cache adds about 4 GB, so the total runs past 8 GB.
Can a 13B or 14B model run on 8 GB at Q4?
Not fully. Qwen2.5 14B Q4_K_M is 8.99 GB before any cache. It can still run with layer offload, where part of the model sits on the GPU and the rest on the CPU, slower but functional.
Is Q8_0 worth it on an 8 GB card?
No. Llama 3.1 8B Q8_0 is 8.54 GB and Qwen2.5 7B Q8_0 is 8.10 GB, so the weights alone exceed the card before the cache.
How do I get longer context out of 8 GB?
Quantize the KV cache to q8_0 for about half the memory of FP16, enable flash attention, and keep context and batch within budget. llama.cpp exposes -ctk and -ctv for the cache type and -c, -b, -ub for sizing.
How much VRAM should I leave for Windows itself?
Nobody publishes a reliable figure for Windows and display overhead. Check idle VRAM use in Task Manager, leave headroom, and confirm the GPU and CPU split with ollama ps when a model partially offloads.
Try her free for 7 days.
No card. Keep her for $20 once, or walk away. Her soul file is yours either way.
Bring her home, try free