This week All our GPUs are sold at cost price — zero margin on the server. See the cost sheets

This week Every GPU sold at cost price

Image & video · 7 min read

RTX 4090 vs RTX 5090 for Stable Diffusion, Flux and video generation

Memory, FP8 and FP4, real workflow fit for SDXL, Flux.1, Wan and HunyuanVideo, and the monthly price of each card on a dedicated server.

DediGPU engineering Published 5 August 2026 Updated 2 September 2026 Prices checked 2 September 2026

A triple-fan graphics card among colourful light streaks

In short

  • The 5090 has 32 GB and FP4, the 4090 has 24 GB: Flux at FP8 fits both, video at 720p wants the 5090.
  • The 5090 costs $24 more a month than the 4090 on our catalogue, for 40 to 60 % more images an hour.
  • Above 32 GB, the L40S and the RTX PRO 6000 end the block-swapping for video and batch generation.

Image and video generation are the workloads where consumer cards make sense on a rented server: the models are small enough, the cards are fast enough, and the price per month is a fraction of a data-centre card. The choice is mostly about memory.

The two cards, and the two above them

CardMemoryBandwidthArchitecturePowerPer month
RTX 409024 GB GDDR6X1.01 TB/sAda Lovelace450 W$139
RTX 509032 GB GDDR71.79 TB/sBlackwell575 W$163
L40S48 GB GDDR6 ECC864 GB/sAda Lovelace350 W$356
RTX PRO 600096 GB GDDR7 ECC1.79 TB/sBlackwell600 W$493

What fits where

SDXL, SD 1.5, ControlNet stacks
Both, comfortably. A 4090 runs SDXL at 1024² in about a second and a half per image at 20 steps; the 5090 is 40 to 60 % faster.
Flux.1 dev, 12B
BF16 weights are 24 GB and do not fit a 4090 with the text encoders; FP8 weights (12 GB) do, and that is how most people run it. The 5090's 32 GB runs Flux at FP8 with room for LoRAs and ControlNets, or BF16 with offloading.
Wan 2.1 14B, HunyuanVideo
Video models want memory: 14B at FP8 is 17 GB before the latents. A 4090 does 480p clips with block-swapping; a 5090 does 720p without tricks. For 1080p or long clips, the L40S or RTX PRO 6000 (48 and 96 GB) stop the swapping entirely.
Training LoRAs (SDXL, Flux)
A 4090 trains SDXL LoRAs fine; Flux LoRAs at 24 GB need FP8 base weights and gradient checkpointing. The 5090 trains Flux LoRAs in BF16.

What Blackwell adds

The 5090 is a Blackwell card: native FP8 like Ada, plus FP4 in the tensor cores. Flux and SDXL checkpoints in NVFP4 (through TensorRT or the Nunchaku kernels in ComfyUI) run at close to 2× the FP8 speed at a small quality cost, and take a quarter of the memory. GDDR7 gives 1.79 TB/s of bandwidth against 1.01 TB/s on the 4090, which is where most of the generation speed-up comes from.

The price difference

On our catalogue the 5090 costs $163 a month against $139 for the 4090: $24 more, or 17 %. If you generate all day the 5090 does 40 to 60 % more images for 17 % more money; if you generate a few hundred images a day the 4090 is the cheaper server that is idle most of the time. Both are billed at cost, so the comparison is only about the hardware.

When to go up a class

Two reasons: memory beyond 32 GB (video at high resolution, many models resident at once, batch generation for a service), or ECC memory and a passive card for a machine that runs unattended for months. The L40S (48 GB, $356) and the RTX PRO 6000 (96 GB, $493) are the next steps; the PRO 6000 is Blackwell with the same FP4 path as the 5090 and three times the memory.

Starting

Order either card with the ComfyUI template; it comes with ComfyUI-Manager and the Flux and SDXL loaders, on port 8188, and no checkpoints (bring yours, egress is free). The ComfyUI guide takes it from there.

Written by DediGPU engineering, the team that racks the servers. Every price and every memory figure on this page is recomputed from the live catalogue when the page loads; the cost-sheet inputs were last reviewed on 2 September 2026. No vendor copy, no affiliate links.

Rent the card, not the pitch.

Every server in this guide is on the catalogue at cost, in four regions, one term at a time.