*[DediGPU](https://dedigpu.com/) — Markdown mirror of [https://dedigpu.com/guides](https://dedigpu.com/guides) · updated 2026-09-02 · index for LLMs: [llms.txt](https://dedigpu.com/llms.txt) · everything: [llms-full.txt](https://dedigpu.com/llms-full.txt)*

Guides

# GPU guides, before you rent.

10 guides written by the team that runs the hardware, with the prices and the memory math recomputed from the catalogue every time the page loads. No vendor copy, no affiliate links.

[Inference How much VRAM do you need to run an LLM? Every size, every precision A single rule, a table for eleven popular models from 8B to 671B at BF16, FP8, INT8 and INT4, and the cheapest dedicated server that fits each one. 2 September 2026 · 9 min read](https://dedigpu.com/guides/how-much-vram-to-run-an-llm)

- [Economics Renting vs buying an H100: the 36-month math, line by line What an H100 costs to own and run over three years, from the same cost sheet that prices our servers, and the break-even against renting at cost. 2 September 2026 · 8 min read](https://dedigpu.com/guides/renting-vs-buying-an-h100)
- [Hardware H100 vs H200 vs B200 vs B300: which one to rent in 2026 Memory, bandwidth, NVLink and price per month of the four NVIDIA data-centre generations on the catalogue, and a plain answer for training, fine-tuning and inference. 2 September 2026 · 8 min read](https://dedigpu.com/guides/h100-vs-h200-vs-b200-vs-b300)
- [Inference Serve Llama 3.3 70B with vLLM on a dedicated GPU server, step by step Pick the server that fits, pre-load the model at order time, check the OpenAI-compatible API, tune context and parallelism, and put a token in front of it. 2 September 2026 · 10 min read](https://dedigpu.com/guides/serve-llama-70b-with-vllm)
- [Training Multi-GPU training: NVLink, InfiniBand, and when one 8× node beats two 4× servers How the interconnect decides training speed: PCIe vs NVLink bandwidth, tensor and pipeline parallelism, InfiniBand between nodes, and how to size a cluster. 2 September 2026 · 8 min read](https://dedigpu.com/guides/multi-gpu-nvlink-infiniband)
- [Image & video RTX 4090 vs RTX 5090 for Stable Diffusion, Flux and video generation Memory, FP8 and FP4, real workflow fit for SDXL, Flux.1, Wan and HunyuanVideo, and the monthly price of each card on a dedicated server. 2 September 2026 · 7 min read](https://dedigpu.com/guides/rtx-4090-vs-rtx-5090-for-diffusion)
- [Economics Dedicated GPU server vs GPU cloud: what the hourly price hides The lines that do not appear on an hourly GPU price: storage that keeps billing, egress, idle hours, pre-emption, shared hosts. And when the cloud is still the right call. 2 September 2026 · 7 min read](https://dedigpu.com/guides/dedicated-gpu-server-vs-gpu-cloud)
- [Fine-tuning Fine-tune an 8B model with QLoRA on a single RTX 4090 Everything that fits in 24 GB: 4-bit base weights, LoRA adapters, a real dataset, the training script, memory numbers, hours, and what it costs on a server billed at cost. 2 September 2026 · 9 min read](https://dedigpu.com/guides/qlora-fine-tune-on-one-rtx-4090)
- [Billing How to pay for a GPU server in crypto, without KYC, and get a refund if you need one The prepaid balance explained: which coin to use and why, fees and confirmation times, what to do when a payment is short, how refunds go back to the address they came from. 2 September 2026 · 6 min read](https://dedigpu.com/guides/pay-for-a-gpu-server-with-crypto)
- [Image & video Run ComfyUI on a rented GPU server: setup, checkpoints, secure access Order a server with the ComfyUI template, reach it safely through an SSH tunnel, bring your checkpoints, add custom nodes, and keep it running for months. 2 September 2026 · 7 min read](https://dedigpu.com/guides/run-comfyui-on-a-rented-gpu-server)

## Rent the card the guide recommends.

Same price in four regions, at cost, one term at a time.

[See the servers](https://dedigpu.com/gpu) [Documentation](https://dedigpu.com/docs)
