A GPU for Stable Diffusion, rented by the week or the month.
GPU Rent Hub rents a dedicated NVIDIA RTX 4090 GPU for Stable Diffusion at $390 a month or $110 a week, paid in Bitcoin, Ethereum, USDT, USDC, Monero, Solana or Litecoin, with no identity verification and no email address on the account.
What does Stable Diffusion actually need from a GPU?
Less VRAM than most people fear, and more raw shader throughput than the spec sheets put in front of you.
Image generation is a denoising loop. The model runs the same forward pass 20 to 50 times per image, so wall-clock speed tracks shader and tensor throughput, while VRAM only decides what fits at all. Stable Diffusion 1.5 runs in about 6 GB. SDXL at 1024 by 1024 runs in 12 GB and is comfortable at 16 GB with a batch of four. Flux-class 12 billion parameter transformers need 24 GB with the weights quantised to FP8, and are happier on 32 GB, which holds the transformer in BF16 next to an FP8 text encoder. A LoRA trained on SDXL wants 24 GB. A LoRA trained on Flux wants 48 GB unless you quantise. Past that point more VRAM buys larger batches, not faster single images.
The arithmetic behind those numbers
Weights cost two bytes per parameter in BF16 and one byte in FP8. A 12B transformer is therefore 24 GB in BF16 and 12 GB in FP8. The text encoder most Flux workflows load is the encoder tower of google/t5-v1_1-xxl, 4.7B parameters of the checkpoint's 11B, so 9.4 GB in BF16 and 4.7 GB in FP8. The SDXL UNet is 2.6B parameters, about 5.2 GB in FP16, which is why 12 GB is genuinely enough for it. Add the VAE and your batch activations and the question is only whether the whole graph stays resident, because offloading to system RAM is where throughput goes to die.
Throughput, not tensor marketing
Diffusion spends its time in the shader and tensor paths, not in the sparse FP8 figures quoted for LLM serving. Of the cards where our catalogue reports FP32 shader throughput, the RTX 4090 is 82.6 TFLOPS, the RTX 6000 Ada 91.1 and the RTX 5090 104.8. Divided by monthly price that is 0.212, 0.198 and 0.194 TFLOPS per dollar. The cheap card is the efficient one.
Which GPU should you rent for SDXL, Flux and ComfyUI?
Sorted by how often we deploy it for image work. The per-GPU-hour column is arithmetic, not a meter: monthly over 720.
| GPU | VRAM | Bandwidth | Monthly | Weekly | Per GPU-hour | In stock |
|---|---|---|---|---|---|---|
| RTX 4090Default pick SDXL at batch 8, Flux in FP8 | 24 GB | 1.01 TB/s | $390/mo | $110 | $0.54 | 265 |
| RTX 5090GDDR7, batch throughput, Flux in BF16 | 32 GB | 1.79 TB/s | $540/mo | $155 | $0.75 | 100 |
| L40SECC, datacenter card, production endpoint | 48 GB | 864 GB/s | $520/mo | $148 | $0.72 | 74 |
| RTX A600048 GB when capacity beats speed | 48 GB | 768 GB/s | $290/mo | $82 | $0.40 | 65 |
| RTX 6000 AdaAda silicon with 48 GB and ECC | 48 GB | 960 GB/s | $460/mo | $130 | $0.64 | 26 |
| RTX 3090SD 1.5 and SDXL on a budget | 24 GB | 936 GB/s | $230/mo | $65 | $0.32 | 186 |
| RTX A5000Cheapest 24 GB with ECC, 230 W | 24 GB | 768 GB/s | $190/mo | $55 | $0.26 | 51 |
| L472 W, low bandwidth, avoid for SDXL | 24 GB | 300 GB/s | $250/mo | $70 | $0.35 | 104 |
RTX 4090 at $390
The default. 24 GB and 1.01 TB/s covers SDXL at batch 8, Flux in FP8 and ControlNet stacks, and it is our largest stock.
RTX 5090 at $540
For batch work. 1.79 TB/s of GDDR7 is 77% more memory bandwidth than the 4090, and 32 GB holds a Flux transformer in BF16.
L40S at $520
For an endpoint other people depend on: ECC memory, 48 GB, a passive datacenter card rated for continuous duty at 350 W.
RTX A6000 at $290
When 48 GB matters more than speed: Flux LoRA training, long node chains, several models resident. Ampere, so slower, at $6.04 per GB.
Can you run ComfyUI on a rented GPU, and who sees the output?
Yes, and nobody here. Both halves of that answer matter, so both are written out.
What ships on the ComfyUI template
Ubuntu 24.04, the NVIDIA driver, ComfyUI, ComfyUI-Manager and the common custom nodes, and no models. You supply your own checkpoints, LoRAs and VAEs. We ship no weights because we will not guess your licences. The interface is not exposed for you either: bind it to localhost and forward the port over SSH, as in the deployment documentation.
What we hold
Host telemetry is five operational numbers and none of them describe your work, listed in full on no-KYC GPU hosting. The part that matters for image work is what we do not hold: no agent in your operating system, no console recording, no copy of your prompts, models or output, and a local NVMe key that is destroyed when the term ends. That is a position about what we keep, not a promise about what is knowable. What that means when someone asks is set out on no-KYC GPU hosting.
The acceptable use line, stated once
Our acceptable use policy is a closed four-item list about networks and hardware rather than a content review, written out on no-KYC GPU hosting. It is not a claim that the law stops applying to you.
Minimum deposit $20 equivalent. How crypto payment works.
Does a second GPU make image generation faster?
Not for one image. For a queue of ten thousand, it is close to linear, and that is the case worth buying.
A single diffusion pass does not split usefully across cards, so two GPUs will not halve the time to your first render. What they do is run independent workers against one queue: a ComfyUI process per card, each holding its own copy of the checkpoint. Throughput scales with card count, latency does not move. Ask for two or more GPUs in one instance and the discount applies automatically.
| RTX 4090 configuration | Discount | Monthly | Weekly | Per GPU-hour |
|---|---|---|---|---|
| 1 card | none | $390 | $110 | $0.54 |
| 2 cardsone host | 5% | $741 | $209 | $0.51 |
| 4 cardsone host | 8% | $1,435 | $405 | $0.50 |
| 8× RTX 4090 node64 vCPU · 376 GB · 2 TB NVMe · 100G | 10% | $2,810 | $790 | $0.49 |
| 8× RTX 5090 node96 vCPU · 496 GB · 4 TB NVMe · 100G | 10% | $3,890 | $1,120 | $0.68 |
| 8× L40S node128 vCPU · 752 GB · 4 TB NVMe · 100G | 10% | $3,740 | $1,070 | $0.65 |
How do you get thousands of renders off the box?
Included bandwidth, and one real constraint on the cheapest card that you should plan around.
20 TB outbound included
Per instance, per month, never throttled. At roughly 1.5 MB for a 1024-pixel PNG that is on the order of 13 million images before a bandwidth line appears.
250 GB NVMe on the 4090
The honest constraint. A Flux checkpoint is about 24 GB and an SDXL one about 7 GB, so a real library fills that drive. The 5090, L40S and A6000 ship 500 GB.
Keep the library between terms
A network volume is NVMe-backed, triple-replicated in the region, attachable to several instances, and it outlives the instance that made it.
Weekly terms for burst work
Seven days on a 4090 is $110, or $0.65 per GPU-hour, and it converts to a monthly term without reprovisioning.
Also included: power, cooling, a 100G uplink with a 10G guarantee, IPv4 and IPv6, unmetered inbound, console access and reinstalls. Snapshot storage is $6 per TB per month.
When is renting a GPU for image generation the wrong call?
Three cases where you should spend your money somewhere else, including nowhere.
If you already own a 12 GB or 16 GB card, SDXL is not a reason to rent: that card runs SDXL at 1024 by 1024 today. Flux in FP8 and LoRA training are the reasons, because they are what 12 GB cannot hold.
And take the cheaper card when it is the right one. For SD 1.5 and SDXL with no training, the RTX 3090 at $230 does the work for $160 a month less than the 4090. If you need 48 GB but not speed, the A6000 at $290 beats the L40S at $520.
Rent from us if
The queue is never empty, you are training LoRAs, you want four workers on one host, you move real volumes of output, or you want to pay in crypto without an identity check.
Do not rent from us if
You need a card for an afternoon, your workload is a dozen images a day, or you want a managed generation API. We rent bare metal with root on it, and nothing above the driver is our problem.
Related sizing
Video models change the answer: GPU for AI video. Training past LoRAs: GPU for fine-tuning. Card-by-card capacity: GPU VRAM guide.
Renting a GPU for Stable Diffusion, asked plainly.
What GPU do I need for Stable Diffusion?
Is an RTX 4090 enough for Flux?
How much does it cost to rent a GPU for image generation?
Can I run ComfyUI on a rented GPU server?
Do you look at the images I generate?
What is actually forbidden on a GPU Rent Hub server?
Do I need an account or an email to rent a GPU for AI art?
Can I train a LoRA on a rented GPU?
More in the general FAQ, term arithmetic in monthly versus hourly GPU rental, capacity in DFW-1, IAD-1 and PDX-1.
A 4090 with ComfyUI on it, in about ninety seconds.
No email, no ID, no card. $390 a month or $110 a week, in any of seven coins, in Dallas, Ashburn or Hillsboro.