The GPU cloud market in 2026 is unrecognizable from 2023. What began as a fragmented landscape of specialized providers has matured into a complex ecosystem spanning hyperscalers, neoclouds, and spot marketplaces. Navigating it requires understanding not just pricing, but the structural differences in how these providers approach GPU compute.
This analysis examines the 10 GPU cloud providers and 54 accelerator instance types in the GPUVerse catalog, 4 hyperscalers, 5 neoclouds, and 1 marketplace, and provides a framework for selecting the right provider for your specific workload. For background on how GPUVerse itself works, see What is GPUVerse?.
All prices in this article are as of August 2026. Per-GPU-hour figures quoted as "from" prices are catalog floor prices, the lowest observed marketplace/neocloud rate for that accelerator. Hyperscaler on-demand pricing typically runs well above these floors (H100 on-demand at $4–7/hr is common). Check the live GPU catalog for current numbers.
Provider Categories
Hyperscalers: AWS, Google Cloud, Azure, Oracle
The major cloud providers offer GPU instances as part of their general compute offerings. This integration provides:
- Advantages: Existing enterprise relationships, comprehensive compliance certifications (SOC2, ISO27001, HIPAA, PCI-DSS; FedRAMP on AWS, Azure, and OCI), integrated networking with other cloud services, consistent billing
- Disadvantages: Premium pricing, quota gating, often constrained by regional availability
| Provider | Spot discount | Compliance highlights | Strengths |
|---|---|---|---|
| AWS | ~70% | SOC2, ISO27001, HIPAA, PCI-DSS, FedRAMP | EFA networking, deepest ecosystem |
| Google Cloud | ~60% | SOC2, ISO27001, HIPAA, PCI-DSS | Only first-party TPU access; GPU quota-gated |
| Azure | ~60% | SOC2, ISO27001, HIPAA, PCI-DSS, FedRAMP | InfiniBand ND-series, OpenAI partnership |
| Oracle (OCI) | ~50% | SOC2, ISO27001, HIPAA, PCI-DSS, FedRAMP | Bare-metal superclusters, cluster RDMA, aggressive GPU pricing for a hyperscaler |
Hyperscalers excel for enterprises with existing cloud contracts, strong compliance requirements, or workloads that need to integrate with other managed services. Expect to pay a premium above the floor prices below for that integration.
Neoclouds: CoreWeave, Lambda, Nebius, Crusoe, RunPod
Neoclouds specialize exclusively in GPU compute, offering competitive pricing and specialized infrastructure:
CoreWeave is Kubernetes-native with InfiniBand fabrics and is consistently among the first to field the latest NVIDIA hardware, including Blackwell fleets. Note that CoreWeave offers no spot or preemptible tier, capacity is on-demand and reserved only.
Lambda provides transparent flat pricing and 1-click clusters, which makes it popular with research and training teams. The trade-off is frequent capacity sellouts on popular GPU types; Lambda also offers no spot tier.
Nebius, headquartered in Amsterdam and spun out of Yandex N.V., offers EU data residency with InfiniBand clusters and strong H100 availability. It has positioned itself as a cost-effective option for EU-based AI workloads. No spot tier.
Crusoe differentiates through clean and stranded-energy computing and reserved cluster capacity, appealing to organizations with environmental commitments. No spot tier.
RunPod targets developers with per-second billing, serverless endpoints, and both community and Secure Cloud tiers. Spot capacity is available at roughly 50% off on-demand, and the serverless offering suits inference with variable traffic.
GPU Marketplace: Vast.ai
Vast.ai aggregates capacity from many hosts, including individual operators, and offers the lowest raw GPU prices in the market. It provides both on-demand and interruptible (bid-based) rentals, interruptible capacity runs roughly 60% below on-demand. The trade-off is host-dependent reliability and more variable performance than a single-operator cloud.
The marketplace model is ideal for:
- Batch processing workloads where interruptions are acceptable
- Development and testing environments
- Organizations with engineering capacity to handle interruptions
While hourly rates appear dramatically lower, interruptible capacity carries operational overhead: handling preemptions, managing checkpointing, and dealing with variable availability can consume engineering time that offsets the price savings. Our guide to GPU spot instances covers how to engineer around this.
Pricing Models Explained
Understanding GPU cloud pricing requires familiarity with three distinct billing models:
On-Demand Pricing
Pay per hour (or per second, on RunPod) with no commitment. Highest per-hour cost but maximum flexibility.
H100 80GB pricing as of August 2026:
- Catalog floor (marketplace/neocloud): from $1.99/hr (H100 SXM), from $2.49/hr (H100 PCIe)
- Hyperscaler on-demand: commonly $4–7/hr
The SXM floor sitting below the PCIe floor looks like a typo but is not, it is marketplace supply dynamics. The same applies to the A100 80GB floor ($1.19/hr) sitting below the A100 40GB floor ($1.29/hr).
Spot/Preemptible Pricing
Discounted rates for interruptible instances. Only 6 of the 10 catalog providers offer a spot or interruptible tier, CoreWeave, Lambda, Nebius, and Crusoe do not.
| Provider | Approx. spot discount | Model |
|---|---|---|
| AWS | ~70% | Spot, 2-minute interruption notice |
| Google Cloud | ~60% | Spot/preemptible VMs |
| Azure | ~60% | Spot VMs |
| Oracle (OCI) | ~50% | Preemptible capacity |
| RunPod | ~50% | Spot pods |
| Vast.ai | ~60% | Interruptible bidding |
Reserved/Committed Use
Long-term commitments (1–3 years) in exchange for significant discounts, typically 40–60% below on-demand rates.
Reserved pricing makes sense for:
- Production workloads with predictable capacity needs
- Workloads where spot interruption is unacceptable
- Organizations with sufficient runway to commit capital
The Real Cost: Dollars Per Token
Hourly pricing is nearly meaningless in isolation. The metric that matters is dollars per useful unit of work, dollars per million tokens served, dollars per training step, dollars per batch processed.
The conversion is simple:
Consider two H100 options at the same hourly price but different real-world throughput (throughput figures illustrative):
| Provider | GPU | $/hr | Tokens/sec | $/M Tokens |
|---|---|---|---|---|
| Provider A | H100 80GB | $3.50 | 850 | $1.14 |
| Provider B | H100 80GB | $3.50 | 1,200 | $0.81 |
The arithmetic: $3.50 / (850 × 3,600) × 10⁶ = $1.14 per million tokens; $3.50 / (1,200 × 3,600) × 10⁶ = $0.81. Even at identical hourly rates, Provider B's cost per token is roughly 29% lower because of its superior throughput. This is why GPUVerse's scoring engine normalizes by throughput, not just list price. The throughput gap between GPU generations matters just as much, see the H100 vs H200 vs B200 comparison for that analysis.
Regional Availability in 2026
GPU availability varies significantly by region, which affects both cost and your ability to secure capacity:
North America (US, Canada): Highest GPU density, most provider options, most competitive pricing. US-East remains the default region for most providers.
European Union: GDPR compliance drives EU-region requirements. Nebius and CoreWeave EU regions offer strong H100 availability. Expect a 15–25% premium over US pricing.
Asia-Pacific: Growing availability from hyperscaler regions and regional providers, but limited H100 supply compared to the US and premium pricing due to limited competition.
Middle East: Emerging market with Oracle and local providers. Limited GPU inventory.
Framework for Provider Selection
Given the complexity of this ecosystem, here is a decision framework:
-
For enterprise workloads with compliance requirements: Start with hyperscalers (AWS, Azure, Google Cloud, OCI) where you have existing relationships and need their certifications.
-
For production AI inference: Evaluate CoreWeave, Lambda, and Nebius for competitive pricing and strong availability, keeping in mind none of the three offers a spot tier.
-
For cost-sensitive batch workloads: Consider spot capacity from the six providers that offer it, or interruptible rentals on Vast.ai, if your workload can handle interruptions.
-
For research and experimentation: The marketplace offers the lowest entry cost, but factor in engineering overhead for managing interruptions.
No single provider offers the best price/performance across all GPU types, regions, and use cases. Organizations running production AI workloads increasingly adopt a multi-cloud strategy, using GPUVerse to select providers based on current pricing and availability.
Conclusion
The GPU cloud market in 2026 offers more options than ever, but also more complexity. Success requires understanding not just pricing, but the full cost profile of each option, including throughput differences, spot reliability, and regional constraints.
Teams that evaluate the full provider landscape systematically will consistently find better infrastructure configurations than those locked into single-provider relationships. GPUVerse exists to make that evaluation continuous and explainable.
FAQ
How many GPU cloud providers does GPUVerse compare?
The GPUVerse catalog covers 10 providers, 4 hyperscalers (AWS, Google Cloud, Azure, OCI), 5 neoclouds (CoreWeave, Lambda, Nebius, Crusoe, RunPod), and 1 marketplace (Vast.ai), across 54 GPU and TPU instance types.
What is the cheapest H100 price in 2026?
As of August 2026, the catalog floor for an H100 SXM is $1.99/hr and for an H100 PCIe is $2.49/hr, both on marketplace/neocloud capacity. Hyperscaler on-demand H100 pricing commonly runs $4–7/hr.
Which GPU clouds offer spot instances?
Six of the ten catalog providers: AWS (~70% discount), Google Cloud (~60%), Azure (~60%), OCI (~50%), RunPod (~50%), and Vast.ai (~60%, via interruptible bidding). CoreWeave, Lambda, Nebius, and Crusoe do not offer spot or preemptible tiers.
Is Vast.ai spot-only?
No. Vast.ai offers both on-demand rentals and interruptible bid-based rentals. Interruptible capacity is the cheaper of the two but can be preempted when outbid.
How should I compare GPU prices across providers?
Convert to dollars per unit of useful work. For inference, that is dollars per million tokens: hourly price divided by (tokens-per-second × 3,600), times 10⁶. Two GPUs at the same hourly price can differ substantially in real cost once throughput is accounted for.




