Image & video · 7 min read
RTX 4090 vs RTX 5090 for Stable Diffusion, Flux and video generation
Memory, FP8 and FP4, real workflow fit for SDXL, Flux.1, Wan and HunyuanVideo, and the monthly price of each card on a dedicated server.

In short
- The 5090 has 32 GB and FP4, the 4090 has 24 GB: Flux at FP8 fits both, video at 720p wants the 5090.
- The 5090 costs $24 more a month than the 4090 on our catalogue, for 40 to 60 % more images an hour.
- Above 32 GB, the L40S and the RTX PRO 6000 end the block-swapping for video and batch generation.
Image and video generation are the workloads where consumer cards make sense on a rented server: the models are small enough, the cards are fast enough, and the price per month is a fraction of a data-centre card. The choice is mostly about memory.
The two cards, and the two above them
| Card | Memory | Bandwidth | Architecture | Power | Per month |
|---|---|---|---|---|---|
| RTX 4090 | 24 GB GDDR6X | 1.01 TB/s | Ada Lovelace | 450 W | $139 |
| RTX 5090 | 32 GB GDDR7 | 1.79 TB/s | Blackwell | 575 W | $163 |
| L40S | 48 GB GDDR6 ECC | 864 GB/s | Ada Lovelace | 350 W | $356 |
| RTX PRO 6000 | 96 GB GDDR7 ECC | 1.79 TB/s | Blackwell | 600 W | $493 |
What fits where
- SDXL, SD 1.5, ControlNet stacks
- Both, comfortably. A 4090 runs SDXL at 1024² in about a second and a half per image at 20 steps; the 5090 is 40 to 60 % faster.
- Flux.1 dev, 12B
- BF16 weights are 24 GB and do not fit a 4090 with the text encoders; FP8 weights (12 GB) do, and that is how most people run it. The 5090's 32 GB runs Flux at FP8 with room for LoRAs and ControlNets, or BF16 with offloading.
- Wan 2.1 14B, HunyuanVideo
- Video models want memory: 14B at FP8 is 17 GB before the latents. A 4090 does 480p clips with block-swapping; a 5090 does 720p without tricks. For 1080p or long clips, the L40S or RTX PRO 6000 (48 and 96 GB) stop the swapping entirely.
- Training LoRAs (SDXL, Flux)
- A 4090 trains SDXL LoRAs fine; Flux LoRAs at 24 GB need FP8 base weights and gradient checkpointing. The 5090 trains Flux LoRAs in BF16.
What Blackwell adds
The 5090 is a Blackwell card: native FP8 like Ada, plus FP4 in the tensor cores. Flux and SDXL checkpoints in NVFP4 (through TensorRT or the Nunchaku kernels in ComfyUI) run at close to 2× the FP8 speed at a small quality cost, and take a quarter of the memory. GDDR7 gives 1.79 TB/s of bandwidth against 1.01 TB/s on the 4090, which is where most of the generation speed-up comes from.
The price difference
On our catalogue the 5090 costs $163 a month against $139 for the 4090: $24 more, or 17 %. If you generate all day the 5090 does 40 to 60 % more images for 17 % more money; if you generate a few hundred images a day the 4090 is the cheaper server that is idle most of the time. Both are billed at cost, so the comparison is only about the hardware.
When to go up a class
Two reasons: memory beyond 32 GB (video at high resolution, many models resident at once, batch generation for a service), or ECC memory and a passive card for a machine that runs unattended for months. The L40S (48 GB, $356) and the RTX PRO 6000 (96 GB, $493) are the next steps; the PRO 6000 is Blackwell with the same FP4 path as the 5090 and three times the memory.
Starting
Order either card with the ComfyUI template; it comes with ComfyUI-Manager and the Flux and SDXL loaders, on port 8188, and no checkpoints (bring yours, egress is free). The ComfyUI guide takes it from there.
Keep reading.
Image & videoRun ComfyUI on a rented GPU server: setup, checkpoints, secure accessOrder a server with the ComfyUI template, reach it safely through an SSH tunnel, bring your checkpoints, add custom nodes, and keep it running for months.2 September 2026 · 7 min read
Fine-tuningFine-tune an 8B model with QLoRA on a single RTX 4090Everything that fits in 24 GB: 4-bit base weights, LoRA adapters, a real dataset, the training script, memory numbers, hours, and what it costs on a server billed at cost.2 September 2026 · 9 min read
EconomicsDedicated GPU server vs GPU cloud: what the hourly price hidesThe lines that do not appear on an hourly GPU price: storage that keeps billing, egress, idle hours, pre-emption, shared hosts. And when the cloud is still the right call.2 September 2026 · 7 min read