How much VRAM do you need to run an LLM locally?
Work out the GPU memory a local LLM needs from its parameter count, quantisation and context length, with worked examples from official docs.
You want to run a language model on your own hardware, maybe to keep client documents off third-party servers or to stop paying per token. The first question is always the same: will it fit on the GPU I have, or the one I'm about to buy? Spec sheets and forum threads give wildly different answers. You can work it out yourself in a few minutes, because GPU memory goes on three things you can estimate: the model's weights, the conversation it's holding (the context), and some overhead.
The weights: parameters times bytes per parameter#
A model's weights are just a very large list of numbers, so the memory they need is the number of parameters multiplied by how many bytes each one takes up.
Hugging Face's guide to optimising LLMs for speed and memory gives the rule of thumb: a model with X billion parameters needs roughly 4 × X GB of VRAM in float32 and roughly 2 × X GB in bfloat16 or float16. Its examples are sobering. Llama-2-70b needs about 140 GB in bfloat16, and Falcon-40b about 80 GB. Very few people have that on one card.
That's why most people running models locally use quantised versions, where each weight is stored in fewer bits. The arithmetic stays the same. You just swap in a smaller number of bytes per parameter:
| Precision | Bytes per parameter (approx.) | 8B model | 32B model |
|---|---|---|---|
| 16-bit (bf16/fp16) | 2 | ~16 GB | ~64 GB |
| 8-bit | 1 | ~8 GB | ~32 GB |
| 4-bit | 0.5 | ~4 GB | ~16 GB |
Treat those as floor figures. Real quantisation formats carry a little extra, because they store scaling information alongside the weights. The llama.cpp quantize README lists measured numbers for Llama 3.1 8B. Its popular Q4_K_M format works out at about 4.89 bits per weight and a 4.58 GiB file. Q8_0 is about 8.5 bits per weight and 7.95 GiB, and F16 is 14.96 GiB. The same README notes that, for llama.cpp, memory and disk requirements are currently the same, so the file size on disk is a decent first estimate of the memory the weights will use.
The same table shows something people often don't expect. On the hardware the llama.cpp maintainers measured, the 4-bit version generated text at about 72 tokens per second against about 29 for F16. So a quantised model can be faster as well as smaller. That won't hold on every setup, so measure on your own hardware.
What quantisation costs you#
Smaller isn't free. Hugging Face's guide says plainly that quantisation trades memory efficiency against accuracy, and in some cases against inference time. In its worked example, a model of just over 15 billion parameters needed 32 GB of VRAM at 16-bit, a bit over 15 GB at 8-bit, and about 9.5 GB at 4-bit. The guide notes that the 8-bit version fits on a consumer GPU like the RTX 4090, and that 4-bit puts it within reach of cards such as the RTX 3090 or T4.
How much quality you lose depends on the model, the format and the task, and the llama.cpp README names perplexity and KL divergence as ways to measure it, which beats guessing. The practical approach is to start with the largest precision that fits, then try the next one down. If the work involves careful extraction from contracts or arithmetic, run your own test set on both before you settle on one.
Official tables beat back-of-envelope maths#
Where a model publisher gives memory figures, use them. Google's Gemma 4 model overview has a table of approximate GPU memory needed just to load each model. For the 31B model it gives 69.9 GB at BF16, 34.9 GB at 8-bit (SFP8) and 17.5 GB at 4-bit (Q4_0), and those figures already include 20% for loading overhead.
Two caveats on that page are worth carrying to any model you look at. First, Google says the estimates only cover the static weights. They don't include memory for supporting software or for the context window, and larger context windows need significantly more VRAM on top. Second, the Gemma 4 26B A4B model is a mixture-of-experts design that activates only about 4 billion parameters for each token, but the page says all 26 billion still have to be loaded into memory. Its 4-bit figure is 14.4 GB. So size a mixture-of-experts model by its total parameter count, not the active count in its name. The active count tells you about the work done per token, not the memory you need.
The part people forget: the context window#
While the model works through your prompt and its reply, it keeps a key-value (KV) cache for every token so far. That cache grows with every token, and long documents or long chats make it big.
Hugging Face's guide works through one example: at a sequence length of 16,000 tokens, the KV cache for its example model needed around 15 GB in float16, about half as much as the model's own weights. The same page shows how architecture changes this. With multi-query attention, where attention heads share keys and values, the same cache drops to under 400 MB. That's why two models with the same parameter count can need very different amounts of memory at long context lengths, and why "it fits" in a quick test can become "out of memory" once someone pastes in a 60-page PDF.
Ollama's tools reflect this. Its context length documentation says that setting a larger context length increases the memory needed to run a model, and that by default it picks the context size from your VRAM: 4k tokens under 24 GiB, 32k tokens from 24 to 48 GiB, and 256k at 48 GiB or more. If you're on a 16 GB card and wondering why the model seems to forget the start of a long document, check that default before blaming the model.
Two settings in the Ollama FAQ also matter for memory. Running parallel requests multiplies the context, and the FAQ states that required RAM scales with the number of parallel requests times the context length. In the other direction, you can quantise the KV cache itself. Setting OLLAMA_KV_CACHE_TYPE to q8_0 uses roughly half the memory of the default f16, with what the FAQ describes as a very small loss in precision.
A worked sizing example#
Say you want a private assistant that answers questions about internal policy documents, used by two or three people, with prompts of a few thousand tokens.
An 8B model at Q4_K_M needs about 4.6 GiB for weights, per the llama.cpp figures. Add the KV cache for a few thousand tokens per request and the runtime's own overhead, and an 8 GB card is tight but workable, while 12 to 16 GB gives you room for longer prompts and a second user. A model around 30B at 4-bit needs 15 to 18 GB for weights alone (Gemma 4 31B is 17.5 GB at Q4_0), so a 24 GB card is the realistic starting point, and long contexts will still push you to quantise the KV cache or shorten prompts.
Whatever the arithmetic says, check what actually happened after loading. In Ollama, ollama ps shows a PROCESSOR column. The FAQ explains that "100% GPU" means the model is entirely in GPU memory, while something like "48%/52% CPU/GPU" means part of it spilled into system RAM. A split load still runs, but check its speed in tokens per second before you decide it's good enough.
You don't need an NVIDIA card#
Ollama's hardware support page lists NVIDIA GPUs with compute capability 5.0 or newer, a range of AMD GPUs through ROCm, and GPU acceleration on Apple devices through Metal. Check the supported-hardware list for your exact card before buying, since support varies by generation and driver version.
Start from the model and context length the job really needs, then pick the card. Do it the other way round and you end up shopping by headline VRAM. So write down the largest document someone will paste in, how many people will use it at once, and the smallest model that passes your own test questions. Turn those into weight and KV cache estimates, leave some headroom for the runtime, and you have a hardware spec you can defend.
Sources
- Optimizing LLMs for Speed and Memory — Hugging Face Transformers documentation (checked 2026-10-10)
- llama.cpp quantize tool README — ggml-org/llama.cpp (GitHub) (checked 2026-10-10)
- Gemma 4 model overview — Google AI for Developers (last updated 8 July 2026) (checked 2026-10-10)
- Context length — Ollama documentation (checked 2026-10-10)
- FAQ — Ollama documentation (checked 2026-10-10)
- Hardware support — Ollama documentation (checked 2026-10-10)
Dealing with this in your own business?
Tell us what you're working on and we'll say whether it's worth automating.