local waifu
Bring her home

Pick your platform

Try her free for 7 days. No card. Keep her? $20 once.

New: Local Waifu now runs on Windows 10 and 11. The installer brings everything she needs, nothing else to set up. Windows may show a SmartScreen prompt the first time: click More info, then Run anyway.

blog

Best Small LLMs for Roleplay on 8 GB of RAM

8 min read
In short

On an 8 GB RAM computer, start with a 2B to 4B instruct model in a 4-bit format and keep the context window modest. A model file that fits on disk still needs memory for the runtime, conversation cache, and operating system.

An 8 GB laptop can run a local language model. That sentence is true, but it hides the decision that matters: which model leaves enough memory for the rest of the computer while still producing dialogue you want to continue reading?

For roleplay, the answer is not always the biggest model you can squeeze into memory. A smaller instruct model with a clear character prompt and a sensible context limit can feel better than a larger model that spends its time swapping data or losing the end of the conversation.

This is a guide to choosing, not a benchmark report. I have not run a controlled comparison of every model on every 8 GB machine, so I will not invent tokens-per-second numbers or call one model universally best.

On 8 GB, the practical range is 2B to 4B

The short version: A 2B to 4B model in a 4-bit format is the sensible starting range when the whole computer has 8 GB of memory.

The number in a model name is a rough count of learned parameters. It is not the same as the file size and it does not tell you the final runtime memory by itself. A 4B model can fit in a compact quantized file, but the runtime still needs space for the model state, the active conversation, and the operating system.

Google’s official Gemma 4 overview lists approximate Q4_0 inference memory of 2.9 GB for Gemma 4 E2B and 4.5 GB for Gemma 4 E4B. Those figures are useful planning references, not a promise that every 8 GB computer will run the models comfortably. Your operating system and other open apps still need memory.

The same page explains that lower bit counts reduce memory cost while also reducing capability. That is the trade. A file that loads is not automatically a file that gives a good conversation.

Gemma 4 E2B is the low-memory starting point

The short version: Gemma 4 E2B is a reasonable first model when memory is tight and you want a current small model with text, image, and audio support in the family.

Google describes Gemma 4 E2B as an edge model with an effective 2.3B parameter count and a 128K context window. The official model documentation is at the Gemma 4 model card. The model card also lists text and image input, with audio support for E2B and E4B.

For pure text roleplay, the multimodal features are not the reason to choose it. The reason is the smaller memory footprint and the availability of an instruction-tuned model. A smaller model gives you more room for the rest of the system and reduces the chance that a long conversation becomes a memory problem.

The limitation is equally plain. A small model has less room for subtle planning, complex relationships, and long scenes. It may repeat a phrase, flatten a character’s voice, or follow the latest instruction too strongly. Good prompting helps, but it does not turn a 2B model into a 30B model.

Gemma 4 E4B is the quality step up, if the machine has room

The short version: Gemma 4 E4B sits near the top of the practical range for 8 GB systems, but its official Q4_0 memory figure already takes about 4.5 GB before you account for the rest of the computer.

Google lists E4B at 4.5B effective parameters and an approximate Q4_0 inference memory requirement of 4.5 GB. That can work on an 8 GB machine when the runtime is efficient and the context is controlled. It can also become uncomfortable when you have a browser, voice tools, image software, or a large conversation open at the same time.

Treat the 128K context number as a model capability, not a recommendation for your laptop. A huge context window needs memory for the attention cache. On a small computer, a shorter context often produces a better experience because the system stays responsive.

If your first generation takes a long time, the answer may not be to abandon local roleplay. Reduce the context, close other apps, and try the smaller model. The boring fix is often the useful one.

Qwen3 4B is a solid alternative for text roleplay

The short version: Qwen3 4B is worth considering when you want a compact text model with a clear published context limit and several quantized formats.

The official Qwen3 4B model card lists 4.0B parameters, native context up to 32,768 tokens, and quantized formats including Q4_K_M, Q5_K_M, Q6_K, and Q8_0. It also documents a larger context mode using YaRN, while warning that the extended setup is not the default choice for ordinary shorter prompts.

That warning is important for roleplay. A long advertised context does not guarantee that the model will remember every scene detail or stay in character. Context is the text the model can see in one request. Memory is the information your companion chooses to retrieve and place there. Those are different systems.

For a small computer, Q4_K_M is the place to begin. Move to a higher precision format only if the machine has enough free memory and you can feel a useful improvement. The exact file size varies by model and release, so check the model card before downloading.

Quantization makes the model fit, but it is not magic

The short version: Quantization reduces the memory used by model weights, while output quality and runtime memory still depend on the model, format, context, and inference tool.

The Q in Q4 or Q5 refers to a lower-precision representation of the weights. You can think of it as storing the model with fewer bits per value. This is why a model that is too large in its original precision can fit in a much smaller file.

The Qwen2.5 3B GGUF card shows multiple quantization options and their approximate file sizes. It also states that the 3B model supports 32,768 tokens of context. The file size is only the starting point, though. The runtime, cache, and operating system need their share.

Do not choose Q2 only because it is the smallest download. A broken or repetitive voice defeats the point of a companion. Start with a usable 4-bit format, then move down only when memory leaves you no other practical option.

Context length is the hidden memory bill

The short version: If a small local model feels fine in a short chat but slows down later, the growing context or cache may be the reason.

Every new turn adds text that the model may need to process. A companion app may also add a system prompt, character details, retrieved memories, and formatting instructions. The visible conversation is not the entire request.

Set a modest context limit first. Increase it only when you have a reason, such as a longer scene that genuinely needs more history. If the model starts making slower replies after many turns, do not assume the model suddenly became worse. The request simply became heavier.

This is also why persistent memory is useful. A memory system can retrieve a relevant fact instead of pasting the entire old chat into every request. I explain the difference in A Bigger Context Window Is Not Memory.

The simple 8 GB decision rule

The short version: Start with Gemma 4 E2B or Qwen3 4B at Q4, keep the context modest, and move up only after the small setup is stable.

Use Gemma 4 E2B when memory is tight and fast startup matters most. Try Gemma 4 E4B when you have enough free memory and want more room for language quality. Try Qwen3 4B when you want a compact text model with documented GGUF options and a native 32K context.

Do not compare models by parameter count alone. Check the actual quantized file, the supported chat template, the context setting, and whether your app has to reserve memory for speech or image generation as well.

A local model is a compromise, but it is a useful one. The right small model leaves enough room for the computer to remain a computer. That matters more than winning a ranking you cannot reproduce on your own hardware.

If you are picking a machine rather than a model, the system requirements page lists what the app itself needs on each platform.

Frequently asked questions

What is the best small LLM for roleplay on 8 GB of RAM?

Start with a 2B to 4B instruct model in a 4-bit format. Gemma 4 E2B, Gemma 4 E4B, and Qwen3 4B are reasonable candidates, but the best result depends on the runtime, context setting, and the rest of your system.

Can a 4B model run on 8 GB of RAM?

It can, especially in a quantized format, but the model does not get all 8 GB. The operating system, inference runtime, conversation cache, and other applications also need memory.

Is Q4 enough for roleplay?

Q4 is a sensible starting point for limited hardware. Higher precision can preserve more quality, but the extra memory may make the computer slower or force the model to use system resources it does not have.

Should I use a 128K context window on an 8 GB computer?

Usually not as a starting point. A large context can consume considerable memory. Begin with a smaller setting and increase it only when your hardware remains responsive.

Does a bigger context window create long-term memory?

No. Context is the text the model can see in one request. Long-term memory is a separate system that stores and retrieves important details across conversations.

Questions people ask

What is the best small LLM for roleplay on 8 GB of RAM?

Start with a 2B to 4B instruct model in a 4-bit format. Gemma 4 E2B, Gemma 4 E4B, and Qwen3 4B are reasonable candidates, but the best result depends on the runtime, context setting, and the rest of your system.

Can a 4B model run on 8 GB of RAM?

It can, especially in a quantized format, but the model does not get all 8 GB. The operating system, inference runtime, conversation cache, and other applications also need memory.

Is Q4 enough for roleplay?

Q4 is a sensible starting point for limited hardware. Higher precision can preserve more quality, but the extra memory may make the computer slower or force the model to use system resources it does not have.

Should I use a 128K context window on an 8 GB computer?

Usually not as a starting point. A large context can consume considerable memory. Begin with a smaller setting and increase it only when your hardware remains responsive.

Does a bigger context window create long-term memory?

No. Context is the text the model can see in one request. Long-term memory is a separate system that stores and retrieves important details across conversations.

Try her free for 7 days.

No card. Keep her for $20 once, or walk away. Her soul file is yours either way.

Bring her home, try free

Back to the blog