As of August 2026, cloud GPU floor prices range from $0.59/hr for a marketplace RTX 4090 to $6.99/hr for a GB200, with the workhorse H100 SXM starting at $1.99/hr and the H200 at $2.89/hr. Those are the headline numbers; the useful version of them, who charges what premium above the floor, when spot pricing applies, and how to convert any hourly rate into cost per token, is the rest of this article.
Every figure here is a catalog floor price, the lowest on-demand rate observed across the 10 providers GPUVerse tracks, as of August 12, 2026. GPU pricing moves weekly; treat this as a calibrated snapshot, and the live catalog as the source of truth.
The August 2026 price index
Floor prices across the GPUVerse catalog, cheapest to most expensive. Throughput figures are tensor TFLOPS with sparsity (dense is half):
| Accelerator | VRAM | Memory bandwidth | FP16 tensor | From $/hr | Tier |
|---|---|---|---|---|---|
| RTX 4090 | 24 GB GDDR6X | 1,008 GB/s | 330 TFLOPS | $0.59 | Budget |
| RTX 5090 | 32 GB GDDR7 | 1,792 GB/s | 419 TFLOPS | $0.69 | Budget |
| L4 | 24 GB GDDR6 | 300 GB/s | 242 TFLOPS | $0.71 | Budget |
| A10 | 24 GB GDDR6 | 600 GB/s | 250 TFLOPS | $0.75 | Budget |
| A100 80GB | 80 GB HBM2e | 2,039 GB/s | 312 TFLOPS | $1.19 | Mid |
| L40S | 48 GB GDDR6 | 864 GB/s | 362 TFLOPS | $1.19 | Mid |
| A100 40GB | 40 GB HBM2e | 1,555 GB/s | 312 TFLOPS | $1.29 | Mid |
| TPU v5e (8-chip slice) | 128 GB HBM | 819 GB/s | 197 TFLOPS | $1.52 | Mid |
| H100 SXM | 80 GB HBM3 | 3,350 GB/s | 1,979 TFLOPS | $1.99 | High |
| TPU v6e Trillium (slice) | 256 GB HBM | 1,640 GB/s | 918 TFLOPS | $2.20 | High |
| TPU v4 (8-chip slice) | 256 GB HBM | 1,200 GB/s | 275 TFLOPS | $2.46 | High |
| H100 PCIe | 80 GB HBM3 | 2,000 GB/s | 1,513 TFLOPS | $2.49 | High |
| H200 | 141 GB HBM3e | 4,800 GB/s | 1,979 TFLOPS | $2.89 | High |
| B200 | 192 GB HBM3e | 8,000 GB/s | 2,250 TFLOPS | $5.49 | Flagship |
| GB200 | 192 GB HBM3e | 8,000 GB/s | 2,250 TFLOPS | $6.99 | Flagship |
Two quirks in that table are real market signal, not typos:
- The A100 80GB floor ($1.19) sits below the 40GB ($1.29). The 80GB variant shipped in far greater volume and now floods the secondary market; the 40GB is scarce enough to command a small premium despite being the worse card.
- The H100 SXM floor ($1.99) sits below the PCIe ($2.49) even though SXM is the better part (3,350 vs. 2,000 GB/s bandwidth, NVLink). Neoclouds and marketplaces are saturated with SXM supply from de-commissioned and over-provisioned training clusters, while PCIe cards ship in scarcer workstation-class inventory. When the better hardware is also cheaper, take the better hardware.
Floor vs. what you'll actually pay
The floor is real and bookable, but it is the marketplace and neocloud floor. Across provider tiers, as of August 2026:
- Marketplace (Vast.ai) sets most of the floors above. Host-dependent reliability, no compliance certifications; you are paid for tolerance of variance. Interruptible bidding cuts prices a further ~60 percent.
- Neoclouds (RunPod, Lambda, CoreWeave, Nebius, Crusoe) run near the floor, typically within 10–50 percent of it, with flat transparent pricing (Lambda), per-second billing (RunPod), or quote-driven reserved clusters (CoreWeave, Nebius, Crusoe) that trade rate for guaranteed capacity.
- Hyperscalers (AWS, Google Cloud, Azure, OCI) price the same silicon at a substantial premium, H100 on-demand commonly lands in the $4–7/hr range, in exchange for compliance depth (HIPAA, PCI-DSS, FedRAMP), global regions, and ecosystem integration. OCI is the consistent outlier, pricing aggressively below its hyperscaler peers. The full trade-off is its own article: hyperscalers vs. neoclouds vs. marketplaces.
Spot and interruptible capacity discounts these rates steeply where offered, roughly 50 percent at RunPod and OCI, 60 percent at Vast.ai, Google Cloud, and Azure, up to 70 percent at AWS. Four providers (CoreWeave, Lambda, Nebius, Crusoe) offer no spot tier at all; their model is reserved capacity. When spot fits your workload's interruption tolerance, it is the single largest discount in the market, our spot instance guide covers the engineering it demands.
Hourly price is an input, not an answer
The number that belongs in your unit economics is dollars per unit of useful work. For inference, that is cost per million tokens:
with your sustained throughput in tokens per second. The formula routinely reverses hourly-price rankings. An H200 at $2.89/hr serving a batched 2,000 tok/s costs $0.40/M tokens; an H100 SXM at $1.99/hr serving 1,650 tok/s costs $0.34/M, but if the workload needs two H100s to match the H200's memory, the H100 pair costs $0.67/M and the "more expensive" H200 wins by 40 percent. Sizing questions like that are exactly what our LLM inference GPU guide walks through, and what the GPUVerse decision engine computes across all 54 accelerators and 10 providers deterministically.
For training, the same logic applies with dollars per training step or per epoch: a B200 at $5.49/hr that halves step time against an H100 at $1.99/hr is not "2.8x the price", per step, the gap collapses to roughly 1.4x, before counting the engineering value of finishing in half the wall-clock time.
Where prices go from here
Two forces are visible in the August 2026 snapshot. Hopper prices continue to soften as Blackwell volume ramps, the H100's floor has fallen far enough that it is now the default value pick of the high tier. And the flagship tier premium (B200 at $5.49, GB200 at $6.99) is holding, because demand for frontier-scale training absorbs supply as fast as it lands. If your workload does not need Blackwell's 192 GB or NVLink-domain scale, the discounted Hopper tier is where 2026's price-performance lives.
FAQ
How much does an H100 cost per hour in 2026?
As of August 2026, H100 SXM starts at $1.99/hr and H100 PCIe at $2.49/hr on marketplaces and neoclouds. Hyperscaler on-demand rates for the same GPU commonly run $4–7/hr; spot pricing can cut 50–70 percent where offered.
How much does an H200 cost per hour?
From $2.89/hr at the catalog floor as of August 2026, with hyperscaler on-demand substantially higher. Its 141 GB of HBM3e often lets one H200 replace two 80 GB GPUs, which is why its cost-per-token economics beat its hourly price.
What is the cheapest cloud GPU?
The RTX 4090 at $0.59/hr (marketplace floor, August 2026) is the cheapest listed accelerator, and for models up to ~14B parameters it is usually the cheapest per token, too, which is the metric that matters.
Why do prices for the same GPU vary so much between clouds?
You are paying for different products around identical silicon: compliance certifications, SLA-backed reliability, cluster networking, global regions, and ecosystem services. A marketplace host offers none of that at $1.99/hr; a FedRAMP-certified hyperscaler region offers all of it at $6/hr. Which premium is worth paying is a workload decision, not a universal one.
Are spot GPUs worth it?
For interruption-tolerant work, training with checkpoints, batch inference, experimentation, usually yes: 50–70 percent savings where offered. For latency-sensitive serving, usually no. Six of the ten providers offer a spot tier; the spot guide covers making preemption survivable.
Should I wait for Blackwell prices to drop?
If your workload fits in 80–141 GB of VRAM, discounted Hopper (H100 from $1.99, H200 from $2.89) is the better buy today. Blackwell's premium is justified where its 192 GB capacity or NVLink-domain scale is the difference between fitting and sharding.




