This week All our GPUs are sold at cost price — zero margin on the server. See the cost sheets

This week Every GPU sold at cost price

Inference · 9 min read

How much VRAM do you need to run an LLM? Every size, every precision

A single rule, a table for eleven popular models from 8B to 671B at BF16, FP8, INT8 and INT4, and the cheapest dedicated server that fits each one.

DediGPU engineering Published 12 August 2026 Updated 2 September 2026 Prices checked 2 September 2026

Macro view of HBM memory stacks next to a GPU die

In short

  • Memory needed ≈ parameters × bytes per parameter × 1.2: a 70B model is 168 GB at BF16, 84 GB at FP8, 42 GB at INT4.
  • Fitting is not fast: leave 30 to 50 % of memory free for the KV cache when you serve more than a few users.
  • The cheapest server that fits each model and precision is computed live from the catalogue in the table below.

The question behind most GPU rentals is “will it fit?”. The answer needs one multiplication, and once you have the number the catalogue tells you what it costs per month. This guide gives the rule, the table and the exceptions.

The rule

Memory needed ≈ parameters × bytes per parameter × 1.2. The bytes depend on precision: 2 for BF16 or FP16, 1 for FP8 or INT8, 0.5 for INT4. The 1.2 covers the KV cache at a normal context length, the activations and the runtime's buffers. A 70B model at BF16 is 70 × 2 × 1.2 = 168 GB; at FP8 it is 84 GB; at INT4, 42 GB.

This is the exact rule our order form applies when you pre-load a model on a serving template: it refuses a combination that does not fit rather than letting you find out at boot.

Eleven models, three precisions, the cheapest server that fits

Prices are per month at cost, from the catalogue, for the smallest server whose total memory covers the need. Mixture-of-experts models count every expert: the whole model sits in memory even if only some experts fire per token.

ModelBF16FP8INT4
Llama 3.1 8B20 GB
1× RTX 4090 $139/mo
10 GB
1× RTX 4090 $139/mo
5 GB
1× RTX 4090 $139/mo
Mistral Small 3.1 24B58 GB
2× RTX 5090 $326/mo
29 GB
1× RTX 5090 $163/mo
15 GB
1× RTX 4090 $139/mo
Gemma 3 27B65 GB
2× RTX A6000 $416/mo
33 GB
2× RTX 4090 $278/mo
17 GB
1× RTX 4090 $139/mo
Qwen3 32B77 GB
2× RTX A6000 $416/mo
39 GB
2× RTX 4090 $278/mo
20 GB
1× RTX 4090 $139/mo
Llama 3.3 70B168 GB
1× MI300X $782/mo
84 GB
1× RTX PRO 6000 $493/mo
42 GB
1× RTX A6000 $208/mo
Qwen2.5 72B173 GB
1× MI300X $782/mo
87 GB
1× RTX PRO 6000 $493/mo
44 GB
1× RTX A6000 $208/mo
Llama 4 Scout · 109B MoE262 GB
2× MI300X $1,564/mo
131 GB
1× MI300X $782/mo
66 GB
2× RTX A6000 $416/mo
Mixtral 8x22B · 141B MoE339 GB
2× MI300X $1,564/mo
170 GB
1× MI300X $782/mo
85 GB
2× RTX A6000 $416/mo
Qwen3 235B-A22B · MoE564 GB
4× MI300X $3,128/mo
282 GB
2× MI300X $1,564/mo
141 GB
1× MI300X $782/mo
Llama 3.1 405B972 GB
8× MI300X $6,256/mo
486 GB
4× MI300X $3,128/mo
243 GB
8× RTX 5090 $1,304/mo
DeepSeek-V3 / R1 · 671B MoE1611 GB
8× B300 $14,096/mo
806 GB
8× MI300X $6,256/mo
403 GB
4× MI300X $3,128/mo

Context length changes the answer

The KV cache grows linearly with the number of tokens in flight: batch size × context length. At 8k tokens and a handful of concurrent requests the 20 % margin is enough. At 128k context or at high concurrency the cache becomes the main consumer. Rule of thumb for Llama-style 70B models: about 320 KB per token at BF16 across all layers, so 128k tokens of context is another 40 GB. Serve long contexts on H200 (141 GB), B200 (180 GB) or MI300X, or spread across more cards with tensor parallelism.

What quantisation costs you

  • FP8 on Ada, Hopper, Blackwell and MI300X is close to lossless for inference and halves memory. It is the default we suggest on those cards.
  • INT8 (SmoothQuant, LLM.int8) works on any card, including Ampere, with a small quality cost.
  • INT4 (AWQ, GPTQ, GGUF Q4) divides memory by four and is fine for chat and most extraction work; it shows on hard reasoning and code. Throughput per card is usually higher, not lower, because the model is memory-bound.

Fitting is not the same as fast

A model that just fits leaves little room for the KV cache, which means small batches and low throughput. If you serve more than a few users, leave 30 to 50 % of memory free. Our catalogue makes that cheap to check: compare a 1× H100 (80 GB, $1,134) with a 1× H200 (141 GB, $1,252): the H200 costs 10 % more and holds 76 % more cache.

What to rent, in short

8B to 32B, one user or a few
1× RTX 4090 or RTX 5090 at INT4/FP8, 1× L40S or RTX PRO 6000 at BF16.
70B class, production
1× H200 or 1× MI300X at FP8 with room for cache; 2× H100 at FP8; 4× H100 at BF16.
100B to 250B MoE
2× H200 or 2× B200 at FP8; 1× B300 for the 235B class at INT4 with cache to spare.
405B and 671B
8× H200 or 8× B200 at FP8; 4× B300 at FP8 for 405B. These are cluster-class servers; see clusters.

Every number above is recomputed from the catalogue when prices change, so the table is never older than the last price review (2 September 2026).

Written by DediGPU engineering, the team that racks the servers. Every price and every memory figure on this page is recomputed from the live catalogue when the page loads; the cost-sheet inputs were last reviewed on 2 September 2026. No vendor copy, no affiliate links.

Rent the card, not the pitch.

Every server in this guide is on the catalogue at cost, in four regions, one term at a time.