Spot instances can reduce GPU compute costs by 50-70%. For batch training workloads that can tolerate interruption, or inference workloads with sufficient redundancy, spot pricing represents an enormous opportunity. Yet many teams who adopt spot instances find the operational complexity erodes, or completely eliminates, the cost savings.
This guide provides a practical framework for using GPU spot instances effectively, with specific strategies for different workload types.
All prices and discount figures in this article are market observations as of August 2026. See the GPU catalog for current floor prices.
How Spot Instances Work
GPU cloud providers maintain pools of spare capacity, GPUs not currently allocated to on-demand or reserved customers. Rather than letting this capacity sit idle, some providers offer it at significant discounts with one critical catch: the capacity can be reclaimed with little notice.
The first thing to know is that not every GPU cloud offers spot at all. Of the ten providers in the GPUVerse catalog, six do:
| Provider | Typical discount (as of Aug 2026) | Interruption model |
|---|---|---|
| AWS | ~70% | 2-minute notice via instance metadata |
| Google Cloud | ~60% | ~30-second preemption notice |
| Azure | ~60% | ~30-second notice via Scheduled Events |
| Oracle Cloud (OCI) | ~50% | Preemptible capacity reclaimed with minimal notice |
| RunPod | ~50% | Short SIGTERM window; per-second billing softens partial runs |
| Vast.ai | ~60% | Interruptible bidding, instance pauses when outbid |
CoreWeave, Lambda, Nebius, and Crusoe do not offer spot or preemptible capacity. Their discount mechanism is different: reserved and committed contracts at negotiated rates, with on-demand as the flexible tier.
No provider publishes preemption rates. Any per-provider percentage you see quoted is anecdotal, actual interruption frequency varies by region, GPU type, time of day, and overall market demand, and can swing from near-zero to a large fraction of instances during capacity crunches. Measure your own rates per provider and region; don't plan around someone else's.
AWS EC2 spot instances get a 2-minute interruption notice but can experience unexpected interruptions during system rebalances. AWS also occasionally terminates spot instances without advance notice in extreme capacity situations. Factor this into your reliability calculations.
Calculating True Spot Savings
The headline savings are impressive. Using illustrative hyperscaler on-demand rates (as of August 2026, hyperscaler on-demand H100s commonly run $4–7/hr):
| GPU | On-Demand/hr (hyperscaler) | Spot/hr | Savings |
|---|---|---|---|
| H100 80GB | $4.50 | $2.25 | 50% |
| A100 80GB | $2.75 | $1.25 | 55% |
Keep the baseline honest: these are hyperscaler on-demand numbers. Neocloud and marketplace on-demand floors are far lower, H100 SXM from $1.99/hr and H200 from $2.89/hr as of August 2026, so always compare spot against your realistic on-demand alternative, not the most expensive one. Sometimes a cheap on-demand provider beats an expensive provider's spot tier; that cross-provider comparison is the core of a multi-cloud GPU strategy.
But true cost savings require accounting for:
- Preemption overhead: Time spent restarting interrupted workloads
- Checkpointing costs: Storage and compute for saving state
- Idle capacity during preemption: Time waiting for replacement instances
- Operational complexity: Engineering time to manage spot lifecycle
A Realistic Example
Consider a training job requiring 1,000 H100-hours. Running on spot at $2.25/hr against the $4.50/hr on-demand baseline above:
Spot scenario with 10% preemption rate:
- Raw compute cost: 1,000 × $2.25 = $2,250
- Preemption overhead (10% extra time): +$225
- Checkpointing storage: +$50
- True cost: $2,525
On-demand scenario at $4.50/hr:
- Raw compute cost: 1,000 × $4.50 = $4,500
- No preemption overhead
- No checkpointing costs
- True cost: $4,500
Even with a 10% preemption rate, spot instances save ~44%. However, if your preemption rate climbs to 30% (common during peak demand periods), the true savings shrink significantly.
Workload Suitability Assessment
Not all workloads are suitable for spot instances. Use this framework:
Suitable for Spot
| Workload Type | Rationale |
|---|---|
| Batch training with checkpoints | Can save progress, resume after interruption |
| Asynchronous inference | Queue-based, stateless inference servers |
| CI/CD training runs | Independent jobs that can be retried |
| Large-scale sweeps | Many parallel jobs, some failures acceptable |
| Image/video processing | Independent frame processing |
Not Suitable for Spot
| Workload Type | Rationale |
|---|---|
| Real-time inference APIs | User-facing, cannot tolerate interruption |
| Training with long iterations | Checkpoint frequency limited by overhead |
| Interactive workloads | User waiting on results |
| Models requiring strict ordering | Preemption breaks reproducibility |
Implementing Spot-Resilient Training
For training workloads that can tolerate preemption, here's an architecture that holds up in production:
Checkpoint-Based Recovery
The example below is complete and runnable as-is (it trains a toy model; swap in your own model, optimizer, and data):
import glob
import os
import signal
import sys
from contextlib import contextmanager
import torch
class CheckpointManager:
def __init__(self, checkpoint_dir, save_frequency=100, keep_last=3):
self.checkpoint_dir = checkpoint_dir
self.save_frequency = save_frequency
self.keep_last = keep_last
self.step = 0
os.makedirs(checkpoint_dir, exist_ok=True)
def save_checkpoint(self, model, optimizer):
"""Save model + optimizer state with an atomic write."""
checkpoint_path = os.path.join(
self.checkpoint_dir, f"checkpoint_{self.step:08d}.pt"
)
temp_path = checkpoint_path + ".tmp"
state = {
"step": self.step,
"model": model.state_dict(),
"optimizer": optimizer.state_dict(),
}
# Atomic write: fully write the file, then rename into place
torch.save(state, temp_path)
os.rename(temp_path, checkpoint_path)
self._cleanup_old_checkpoints()
def _cleanup_old_checkpoints(self):
"""Keep only the most recent `keep_last` checkpoints."""
checkpoints = sorted(
glob.glob(os.path.join(self.checkpoint_dir, "checkpoint_*.pt"))
)
for old in checkpoints[: -self.keep_last]:
os.remove(old)
def load_latest(self, model, optimizer):
"""Restore the newest checkpoint. Returns the step to resume from."""
checkpoints = sorted(
glob.glob(os.path.join(self.checkpoint_dir, "checkpoint_*.pt"))
)
if not checkpoints:
return 0
state = torch.load(checkpoints[-1], map_location="cpu")
model.load_state_dict(state["model"])
optimizer.load_state_dict(state["optimizer"])
self.step = state["step"]
return self.step
def should_save(self):
self.step += 1
return self.step % self.save_frequency == 0
@contextmanager
def spot_preemption_handler(manager, model, optimizer):
"""Save a final checkpoint if the provider sends SIGTERM before reclaiming."""
def handler(signum, frame):
manager.save_checkpoint(model, optimizer)
sys.exit(0)
original = signal.signal(signal.SIGTERM, handler)
try:
yield
finally:
signal.signal(signal.SIGTERM, original)
if __name__ == "__main__":
# Toy training loop: replace with your model, optimizer, and data
model = torch.nn.Linear(1024, 1024)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4)
manager = CheckpointManager("/tmp/checkpoints", save_frequency=100)
start_step = manager.load_latest(model, optimizer)
print(f"Resuming from step {start_step}")
with spot_preemption_handler(manager, model, optimizer):
while manager.step < 10_000:
inputs = torch.randn(32, 1024)
loss = model(inputs).pow(2).mean()
optimizer.zero_grad()
loss.backward()
optimizer.step()
if manager.should_save():
manager.save_checkpoint(model, optimizer)Note that the handler receives the model and optimizer explicitly, signal handlers that reach for globals are a classic source of "checkpoint saved nothing" surprises.
Spot Instance Health Monitoring
On AWS, poll the instance metadata service for interruption notices. Two details matter: IMDSv2 requires a session token (fetched via PUT), and the endpoint returns HTTP 404 with an empty body when no interruption is scheduled, never compare the response against the string "null".
#!/bin/bash
# AWS spot interruption watcher (IMDSv2)
while true; do
TOKEN=$(curl -sS -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 60")
HTTP_CODE=$(curl -s -o /tmp/spot-action -w "%{http_code}" \
-H "X-aws-ec2-metadata-token: $TOKEN" \
"http://169.254.169.254/latest/meta-data/spot/instance-action")
# 404 (empty body) means no interruption is scheduled
if [ "$HTTP_CODE" = "200" ] && [ -s /tmp/spot-action ]; then
echo "Interruption scheduled: $(cat /tmp/spot-action)"
# Stop accepting new work
kubectl scale deployment inference-server --replicas=0
# Save current state (your checkpoint hook)
save_checkpoint
exit 0
fi
sleep 5
doneOn Google Cloud and Azure the equivalent signals are the preemption metadata endpoint and Scheduled Events, respectively; the pattern, poll, treat "not found" as normal, drain and checkpoint on notice, is identical. Wire the notice into your alerting so a termination on a busy instance pages someone, as covered in our production GPU monitoring guide.
Spot Instance Strategies by Provider
Hyperscalers (AWS, Google Cloud, Azure, OCI)
The deepest spot pools and the largest discounts (~50–70% as of August 2026):
- Strategy: Default choice for large-scale interruption-tolerant training
- Bidding: There is none on AWS anymore, user bidding was retired in 2017; you simply pay the current spot price, capped at on-demand. GCP and Azure similarly use provider-set spot prices
- Allocation: Use capacity-optimized allocation (AWS) and diversify across instance types and zones to reduce interruptions
- Fallback: Mixed fleets with an on-demand baseline for critical progress
RunPod
- Strategy: Good for burst batch work and experimentation; ~50% below on-demand
- Billing: Per-second billing means a preempted partial hour costs only what you used
- Tiers: Community cloud is cheapest; Secure Cloud trades some discount for more predictable infrastructure
Vast.ai Interruptible
Vast.ai operates as a marketplace, and its interruptible tier is the one place in this list where bidding is real: your bid competes against other users for a host's capacity, and your instance pauses when outbid.
- Strategy: Best for embarrassingly parallel workloads
- Bidding: Your bid price directly controls stability, a common starting point is around 70% of the host's on-demand price, then adjust based on observed interruptions. Bid low for maximum savings on retry-friendly jobs; bid near on-demand when an interruption would cost you real progress
- Reliability: Variable, depends on the specific host machine as well as your bid
- Best practices: Run multiple smaller instances rather than one large instance; note Vast.ai also offers regular on-demand instances when you need stability
No spot at all: CoreWeave, Lambda, Nebius, Crusoe
These neoclouds don't sell interruptible capacity. If your workload fits their strengths (Kubernetes-native infrastructure, flat pricing, EU residency, clean energy), the cost lever is reserved or committed contracts rather than spot, often competitive with hyperscaler spot pricing once you account for zero preemption overhead. Our 2026 GPU cloud comparison covers where each fits.
Building a Spot Instance Architecture
A production-ready spot architecture for distributed training:
Key principles:
- Scheduler runs on on-demand instances (low GPU requirements, high reliability needs)
- Workers run on spot instances (GPU-intensive, interruption-tolerant)
- Checkpoint to persistent storage (S3, GCS, or equivalent)
- Implement graceful degradation when workers are preempted
Conclusion
GPU spot instances offer real, substantial savings for the right workload types. The key to capturing those savings is honest assessment of your workload's spot-suitability and investment in the operational infrastructure to handle preemption gracefully.
The organizations that use spot instances most effectively treat preemption as a normal operational event, not an edge case. Build your training infrastructure to expect interruptions, checkpoint frequently, and design your recovery procedures to be fast and automated.
When implemented well, spot instances can reduce your GPU compute costs by 50% or more without impacting the reliability of your training pipeline. When implemented poorly, they'll cost more in engineering time than they save in compute costs.
The GPUVerse Discover engine incorporates spot pricing and spot availability per provider into its recommendations, including the fact that four of the ten catalog providers don't offer spot at all. Understanding the underlying mechanics helps you interpret those recommendations and make informed decisions about your spot strategy.
FAQ
Which GPU cloud providers offer spot instances?
As of August 2026, six of the ten providers in the GPUVerse catalog: AWS (~70% discount), Google Cloud (~60%), Azure (~60%), OCI (~50%), RunPod (~50%), and Vast.ai (~60%, via interruptible bidding). CoreWeave, Lambda, Nebius, and Crusoe do not offer spot or preemptible capacity, their discounts come from reserved contracts.
How much do spot instances really save after preemption overhead?
Less than the headline discount, but still substantial. In our worked example, 1,000 H100-hours at $2.25/hr spot vs a $4.50/hr hyperscaler on-demand baseline, a 10% preemption rate plus checkpoint storage still leaves ~44% net savings. At a 30% preemption rate the savings shrink significantly, so measure your actual interruption frequency.
How much notice do I get before a spot instance is terminated?
AWS gives a 2-minute notice via instance metadata; Google Cloud and Azure give roughly 30 seconds via their preemption signal and Scheduled Events. OCI preemptible capacity and marketplace instances can be reclaimed with minimal notice. Design checkpointing so that even the shortest window is enough to save state.
How do I detect an AWS spot interruption in code?
Poll the IMDSv2 spot/instance-action endpoint: PUT a session token first, then GET the endpoint with the token header. HTTP 404 (empty body) means no interruption is scheduled; a 200 with a JSON body means termination is coming, drain and checkpoint immediately. Never compare the response body to the string "null".
Does bidding still exist for spot instances?
Mostly no. AWS retired user bidding in 2017, you pay the going spot price, capped at on-demand, and GCP and Azure set spot prices themselves. The exception is Vast.ai's interruptible tier, a genuine bidding market where a higher bid means fewer interruptions; starting around 70% of the host's on-demand price is a reasonable opening bid.
Are spot instances suitable for production inference?
Not for user-facing, real-time APIs, an interruption becomes a user-visible outage. Spot works for asynchronous, queue-based inference with redundancy, and for batch processing. Keep latency-sensitive serving on on-demand or reserved capacity and use spot for the interruption-tolerant tail of your fleet.



