Skip to content
GPUVerse Blog
Comparisons

GPU Cloud Pricing in 2026: What Every Accelerator Costs

Floor prices for 15 GPU and TPU types across 10 clouds as of August 2026, from $0.59/hr RTX 4090s to $6.99/hr GB200, plus the cost-per-token math that turns hourly rates into real unit economics.

GPUVerse Team
7 min read
Share
NVIDIA B200 Blackwell GPU, the current flagship tier of cloud GPU pricing

As of August 2026, cloud GPU floor prices range from $0.59/hr for a marketplace RTX 4090 to $6.99/hr for a GB200, with the workhorse H100 SXM starting at $1.99/hr and the H200 at $2.89/hr. Those are the headline numbers; the useful version of them, who charges what premium above the floor, when spot pricing applies, and how to convert any hourly rate into cost per token, is the rest of this article.

Pricing decays fast

Every figure here is a catalog floor price, the lowest on-demand rate observed across the 10 providers GPUVerse tracks, as of August 12, 2026. GPU pricing moves weekly; treat this as a calibrated snapshot, and the live catalog as the source of truth.

The August 2026 price index

Floor prices across the GPUVerse catalog, cheapest to most expensive. Throughput figures are tensor TFLOPS with sparsity (dense is half):

AcceleratorVRAMMemory bandwidthFP16 tensorFrom $/hrTier
RTX 409024 GB GDDR6X1,008 GB/s330 TFLOPS$0.59Budget
RTX 509032 GB GDDR71,792 GB/s419 TFLOPS$0.69Budget
L424 GB GDDR6300 GB/s242 TFLOPS$0.71Budget
A1024 GB GDDR6600 GB/s250 TFLOPS$0.75Budget
A100 80GB80 GB HBM2e2,039 GB/s312 TFLOPS$1.19Mid
L40S48 GB GDDR6864 GB/s362 TFLOPS$1.19Mid
A100 40GB40 GB HBM2e1,555 GB/s312 TFLOPS$1.29Mid
TPU v5e (8-chip slice)128 GB HBM819 GB/s197 TFLOPS$1.52Mid
H100 SXM80 GB HBM33,350 GB/s1,979 TFLOPS$1.99High
TPU v6e Trillium (slice)256 GB HBM1,640 GB/s918 TFLOPS$2.20High
TPU v4 (8-chip slice)256 GB HBM1,200 GB/s275 TFLOPS$2.46High
H100 PCIe80 GB HBM32,000 GB/s1,513 TFLOPS$2.49High
H200141 GB HBM3e4,800 GB/s1,979 TFLOPS$2.89High
B200192 GB HBM3e8,000 GB/s2,250 TFLOPS$5.49Flagship
GB200192 GB HBM3e8,000 GB/s2,250 TFLOPS$6.99Flagship

Two quirks in that table are real market signal, not typos:

  • The A100 80GB floor ($1.19) sits below the 40GB ($1.29). The 80GB variant shipped in far greater volume and now floods the secondary market; the 40GB is scarce enough to command a small premium despite being the worse card.
  • The H100 SXM floor ($1.99) sits below the PCIe ($2.49) even though SXM is the better part (3,350 vs. 2,000 GB/s bandwidth, NVLink). Neoclouds and marketplaces are saturated with SXM supply from de-commissioned and over-provisioned training clusters, while PCIe cards ship in scarcer workstation-class inventory. When the better hardware is also cheaper, take the better hardware.

Floor vs. what you'll actually pay

The floor is real and bookable, but it is the marketplace and neocloud floor. Across provider tiers, as of August 2026:

  • Marketplace (Vast.ai) sets most of the floors above. Host-dependent reliability, no compliance certifications; you are paid for tolerance of variance. Interruptible bidding cuts prices a further ~60 percent.
  • Neoclouds (RunPod, Lambda, CoreWeave, Nebius, Crusoe) run near the floor, typically within 10–50 percent of it, with flat transparent pricing (Lambda), per-second billing (RunPod), or quote-driven reserved clusters (CoreWeave, Nebius, Crusoe) that trade rate for guaranteed capacity.
  • Hyperscalers (AWS, Google Cloud, Azure, OCI) price the same silicon at a substantial premium, H100 on-demand commonly lands in the $4–7/hr range, in exchange for compliance depth (HIPAA, PCI-DSS, FedRAMP), global regions, and ecosystem integration. OCI is the consistent outlier, pricing aggressively below its hyperscaler peers. The full trade-off is its own article: hyperscalers vs. neoclouds vs. marketplaces.

Spot and interruptible capacity discounts these rates steeply where offered, roughly 50 percent at RunPod and OCI, 60 percent at Vast.ai, Google Cloud, and Azure, up to 70 percent at AWS. Four providers (CoreWeave, Lambda, Nebius, Crusoe) offer no spot tier at all; their model is reserved capacity. When spot fits your workload's interruption tolerance, it is the single largest discount in the market, our spot instance guide covers the engineering it demands.

Hourly price is an input, not an answer

The number that belongs in your unit economics is dollars per unit of useful work. For inference, that is cost per million tokens:

C1M tokens=phourT⋅3600×106C_{\text{1M tokens}} = \frac{p_{\text{hour}}}{T \cdot 3600} \times 10^{6}

with TT your sustained throughput in tokens per second. The formula routinely reverses hourly-price rankings. An H200 at $2.89/hr serving a batched 2,000 tok/s costs $0.40/M tokens; an H100 SXM at $1.99/hr serving 1,650 tok/s costs $0.34/M, but if the workload needs two H100s to match the H200's memory, the H100 pair costs $0.67/M and the "more expensive" H200 wins by 40 percent. Sizing questions like that are exactly what our LLM inference GPU guide walks through, and what the GPUVerse decision engine computes across all 54 accelerators and 10 providers deterministically.

For training, the same logic applies with dollars per training step or per epoch: a B200 at $5.49/hr that halves step time against an H100 at $1.99/hr is not "2.8x the price", per step, the gap collapses to roughly 1.4x, before counting the engineering value of finishing in half the wall-clock time.

Where prices go from here

Two forces are visible in the August 2026 snapshot. Hopper prices continue to soften as Blackwell volume ramps, the H100's floor has fallen far enough that it is now the default value pick of the high tier. And the flagship tier premium (B200 at $5.49, GB200 at $6.99) is holding, because demand for frontier-scale training absorbs supply as fast as it lands. If your workload does not need Blackwell's 192 GB or NVLink-domain scale, the discounted Hopper tier is where 2026's price-performance lives.

FAQ

How much does an H100 cost per hour in 2026?

As of August 2026, H100 SXM starts at $1.99/hr and H100 PCIe at $2.49/hr on marketplaces and neoclouds. Hyperscaler on-demand rates for the same GPU commonly run $4–7/hr; spot pricing can cut 50–70 percent where offered.

How much does an H200 cost per hour?

From $2.89/hr at the catalog floor as of August 2026, with hyperscaler on-demand substantially higher. Its 141 GB of HBM3e often lets one H200 replace two 80 GB GPUs, which is why its cost-per-token economics beat its hourly price.

What is the cheapest cloud GPU?

The RTX 4090 at $0.59/hr (marketplace floor, August 2026) is the cheapest listed accelerator, and for models up to ~14B parameters it is usually the cheapest per token, too, which is the metric that matters.

Why do prices for the same GPU vary so much between clouds?

You are paying for different products around identical silicon: compliance certifications, SLA-backed reliability, cluster networking, global regions, and ecosystem services. A marketplace host offers none of that at $1.99/hr; a FedRAMP-certified hyperscaler region offers all of it at $6/hr. Which premium is worth paying is a workload decision, not a universal one.

Are spot GPUs worth it?

For interruption-tolerant work, training with checkpoints, batch inference, experimentation, usually yes: 50–70 percent savings where offered. For latency-sensitive serving, usually no. Six of the ten providers offer a spot tier; the spot guide covers making preemption survivable.

Should I wait for Blackwell prices to drop?

If your workload fits in 80–141 GB of VRAM, discounted Hopper (H100 from $1.99, H200 from $2.89) is the better buy today. Blackwell's premium is justified where its 192 GB capacity or NVLink-domain scale is the difference between fitting and sharding.

GPUVerse Team

GPUVerse Team

AI Infrastructure Engineering

The team building the intelligence layer above every cloud. We write about what we learn operating GPU infrastructure at scale.

Keep reading