Inference · 9 min read
How much VRAM do you need to run an LLM? Every size, every precision
A single rule, a table for eleven popular models from 8B to 671B at BF16, FP8, INT8 and INT4, and the cheapest dedicated server that fits each one.

In short
- Memory needed ≈ parameters × bytes per parameter × 1.2: a 70B model is 168 GB at BF16, 84 GB at FP8, 42 GB at INT4.
- Fitting is not fast: leave 30 to 50 % of memory free for the KV cache when you serve more than a few users.
- The cheapest server that fits each model and precision is computed live from the catalogue in the table below.
The question behind most GPU rentals is “will it fit?”. The answer needs one multiplication, and once you have the number the catalogue tells you what it costs per month. This guide gives the rule, the table and the exceptions.
The rule
Memory needed ≈ parameters × bytes per parameter × 1.2. The bytes depend on precision: 2 for BF16 or FP16, 1 for FP8 or INT8, 0.5 for INT4. The 1.2 covers the KV cache at a normal context length, the activations and the runtime's buffers. A 70B model at BF16 is 70 × 2 × 1.2 = 168 GB; at FP8 it is 84 GB; at INT4, 42 GB.
Eleven models, three precisions, the cheapest server that fits
Prices are per month at cost, from the catalogue, for the smallest server whose total memory covers the need. Mixture-of-experts models count every expert: the whole model sits in memory even if only some experts fire per token.
| Model | BF16 | FP8 | INT4 |
|---|---|---|---|
| Llama 3.1 8B | 20 GB 1× RTX 4090 $139/mo | 10 GB 1× RTX 4090 $139/mo | 5 GB 1× RTX 4090 $139/mo |
| Mistral Small 3.1 24B | 58 GB 2× RTX 5090 $326/mo | 29 GB 1× RTX 5090 $163/mo | 15 GB 1× RTX 4090 $139/mo |
| Gemma 3 27B | 65 GB 2× RTX A6000 $416/mo | 33 GB 2× RTX 4090 $278/mo | 17 GB 1× RTX 4090 $139/mo |
| Qwen3 32B | 77 GB 2× RTX A6000 $416/mo | 39 GB 2× RTX 4090 $278/mo | 20 GB 1× RTX 4090 $139/mo |
| Llama 3.3 70B | 168 GB 1× MI300X $782/mo | 84 GB 1× RTX PRO 6000 $493/mo | 42 GB 1× RTX A6000 $208/mo |
| Qwen2.5 72B | 173 GB 1× MI300X $782/mo | 87 GB 1× RTX PRO 6000 $493/mo | 44 GB 1× RTX A6000 $208/mo |
| Llama 4 Scout · 109B MoE | 262 GB 2× MI300X $1,564/mo | 131 GB 1× MI300X $782/mo | 66 GB 2× RTX A6000 $416/mo |
| Mixtral 8x22B · 141B MoE | 339 GB 2× MI300X $1,564/mo | 170 GB 1× MI300X $782/mo | 85 GB 2× RTX A6000 $416/mo |
| Qwen3 235B-A22B · MoE | 564 GB 4× MI300X $3,128/mo | 282 GB 2× MI300X $1,564/mo | 141 GB 1× MI300X $782/mo |
| Llama 3.1 405B | 972 GB 8× MI300X $6,256/mo | 486 GB 4× MI300X $3,128/mo | 243 GB 8× RTX 5090 $1,304/mo |
| DeepSeek-V3 / R1 · 671B MoE | 1611 GB 8× B300 $14,096/mo | 806 GB 8× MI300X $6,256/mo | 403 GB 4× MI300X $3,128/mo |
Context length changes the answer
The KV cache grows linearly with the number of tokens in flight: batch size × context length. At 8k tokens and a handful of concurrent requests the 20 % margin is enough. At 128k context or at high concurrency the cache becomes the main consumer. Rule of thumb for Llama-style 70B models: about 320 KB per token at BF16 across all layers, so 128k tokens of context is another 40 GB. Serve long contexts on H200 (141 GB), B200 (180 GB) or MI300X, or spread across more cards with tensor parallelism.
What quantisation costs you
- FP8 on Ada, Hopper, Blackwell and MI300X is close to lossless for inference and halves memory. It is the default we suggest on those cards.
- INT8 (SmoothQuant, LLM.int8) works on any card, including Ampere, with a small quality cost.
- INT4 (AWQ, GPTQ, GGUF Q4) divides memory by four and is fine for chat and most extraction work; it shows on hard reasoning and code. Throughput per card is usually higher, not lower, because the model is memory-bound.
Fitting is not the same as fast
A model that just fits leaves little room for the KV cache, which means small batches and low throughput. If you serve more than a few users, leave 30 to 50 % of memory free. Our catalogue makes that cheap to check: compare a 1× H100 (80 GB, $1,134) with a 1× H200 (141 GB, $1,252): the H200 costs 10 % more and holds 76 % more cache.
What to rent, in short
- 8B to 32B, one user or a few
- 1× RTX 4090 or RTX 5090 at INT4/FP8, 1× L40S or RTX PRO 6000 at BF16.
- 70B class, production
- 1× H200 or 1× MI300X at FP8 with room for cache; 2× H100 at FP8; 4× H100 at BF16.
- 100B to 250B MoE
- 2× H200 or 2× B200 at FP8; 1× B300 for the 235B class at INT4 with cache to spare.
- 405B and 671B
- 8× H200 or 8× B200 at FP8; 4× B300 at FP8 for 405B. These are cluster-class servers; see clusters.
Every number above is recomputed from the catalogue when prices change, so the table is never older than the last price review (2 September 2026).
Keep reading.
InferenceServe Llama 3.3 70B with vLLM on a dedicated GPU server, step by stepPick the server that fits, pre-load the model at order time, check the OpenAI-compatible API, tune context and parallelism, and put a token in front of it.2 September 2026 · 10 min read
HardwareH100 vs H200 vs B200 vs B300: which one to rent in 2026Memory, bandwidth, NVLink and price per month of the four NVIDIA data-centre generations on the catalogue, and a plain answer for training, fine-tuning and inference.2 September 2026 · 8 min read
EconomicsRenting vs buying an H100: the 36-month math, line by lineWhat an H100 costs to own and run over three years, from the same cost sheet that prices our servers, and the break-even against renting at cost.2 September 2026 · 8 min read