Choosing a GPU for LLM inference comes down to three questions, answered in order: does the model fit in memory, is the memory fast enough for your latency target, and which fitting option costs the least per token? Hourly price, the number everyone compares first, is the last input, not the first.
This guide walks the method with real arithmetic: the VRAM formula, a fit table across the accelerators in the GPUVerse catalog, and the cost-per-token comparison that actually decides the bill.
Hourly figures are catalog floor prices, the lowest on-demand rate across the 10 providers GPUVerse tracks, as of August 2026. Hyperscaler on-demand rates run substantially above these floors; see the full pricing breakdown.
Step 1: The VRAM math
An inference deployment needs memory for two things: the model weights and the KV cache, plus 10–20 percent overhead for activations, CUDA context, and the serving framework.
Weights are parameter count times bytes per parameter:
where is 2 bytes for FP16/BF16, 1 byte for FP8/INT8, and 0.5 bytes for 4-bit quantization. So a 70B-parameter model needs roughly 140 GB at FP16, 70 GB at FP8, and 35 GB at INT4.
KV cache grows with context length and concurrency. Per token of context, per sequence:
For Llama 3.1 70B (80 layers, 8 KV heads via grouped-query attention, head dimension 128, FP16 cache) that is 2 × 80 × 8 × 128 × 2 bytes ≈ 0.31 MiB per token, about 10 GB per sequence at a 32k context, and 4x that at 128k. Multiply by your target concurrent sequences. This is why a model that "fits" on paper can still OOM under load, and why memory-per-dollar matters as much as compute-per-dollar for serving.
Step 2: What fits where
Applying that math to current hardware (full specs on the GPU catalog; TFLOPS figures below are tensor throughput with sparsity, dense is half):
| Model size | Precision | Weights | Fits comfortably on | Floor price, Aug 2026 |
|---|---|---|---|---|
| 7–8B | FP16 | ~16 GB | RTX 4090 / L4 / A10 (24 GB), RTX 5090 (32 GB) | from $0.59/hr |
| 7–8B | INT4 | ~4 GB | Anything in the catalog | from $0.59/hr |
| 13–14B | FP16 | ~28 GB | RTX 5090 (32 GB), L40S (48 GB) | from $0.69/hr |
| 30–34B | INT4 | ~17 GB | RTX 4090 / L4 (24 GB, tight), L40S (48 GB) | from $0.59/hr |
| 70B | INT4 | ~35 GB | L40S (48 GB, tight), A100/H100 80 GB | from $1.19/hr |
| 70B | FP8 | ~70 GB | H200 (141 GB) single-GPU, or 2× A100/H100 80 GB | from $2.89/hr |
| 70B | FP16 | ~140 GB | 2× H100/A100 80 GB, 2× H200 | from $2.38/hr (2× A100 80GB) |
| 405B | FP8 | ~405 GB | 4× H200 (564 GB), 8× H100, B200/GB200 clusters | multi-GPU territory |
Three practical notes on the table:
- "Fits" means weights plus KV headroom. A 70B FP8 model on a single H100 80 GB leaves ~2 GB for cache after overhead, it loads, then falls over under real traffic. The H200's 141 GB exists almost precisely for this case: same compute as the H100 SXM, 76 percent more memory, 43 percent more bandwidth. The full comparison is in H100 vs. H200.
- Inference is usually bandwidth-bound. At batch sizes typical for serving, token generation speed tracks memory bandwidth more closely than TFLOPS. That is why an H200 (4.8 TB/s) out-serves an H100 SXM (3.35 TB/s) at identical compute specs, and why the L4 (300 GB/s) is a batch-and-budget card, not a latency card.
- Quantization is the biggest lever you control. Moving 70B from FP16 to FP8 halves the memory footprint and roughly halves the hardware cost with minimal quality loss on most workloads; INT4 halves it again with a real but often acceptable quality trade-off. Evaluate on your task before committing.
Step 3: Compare on cost per million tokens
Once two or three configurations fit, rank them the way the GPUVerse engine does, dollars per unit of useful work, not dollars per hour:
where is the effective hourly price and is your measured sustained throughput in tokens per second. Two worked examples with illustrative throughputs (yours will vary with batch size, context, and serving stack, measure before committing):
Serving Llama 3.1 70B at FP8. Option A: 2× H100 SXM at the $1.99/hr floor = $3.98/hr, serving a batched ~1,650 tok/s → $0.67 per million tokens. Option B: a single H200 at $2.89/hr serving ~2,000 tok/s → $0.40 per million tokens. The single bigger-memory GPU wins on cost and removes a layer of tensor-parallel complexity, a result an hourly-price comparison ($3.98 vs. $2.89) understates and a per-GPU comparison ($1.99 vs. $2.89) gets exactly backwards.
Serving an 8B model. An RTX 4090 at $0.59/hr pushing a batched ~2,500 tok/s → $0.066 per million tokens. This is why small-model workloads almost never belong on Hopper-class hardware: an H100 would need to serve ~10,000 tok/s just to match that unit cost.
Step 4: Then, and only then, pick the provider
The same GPU spans a wide price and reliability range across the 10 providers, marketplace floors, neocloud flat rates, and hyperscaler premiums, with different compliance certifications and spot/interruptible options at each tier. Which tier fits is a workload question, dev experiments tolerate marketplace interruptions; a HIPAA-bound API does not, and we break that choice down in hyperscalers vs. neoclouds vs. marketplaces. Or describe the workload to GPUVerse and let the engine run this whole method deterministically across every provider at once.
FAQ
What GPU do I need to run a 70B model?
At 4-bit quantization, ~35 GB of weights fit on a single 48–80 GB card (L40S tightly, A100/H100 comfortably). At FP8 you want a single H200 (141 GB) or two 80 GB GPUs. At FP16, plan on two 80 GB-class GPUs minimum. Add KV-cache headroom for your context length and concurrency in every case.
How much VRAM does Llama 3.1 70B need?
Roughly 140 GB at FP16, 70 GB at FP8, or 35 GB at INT4 for weights alone, plus KV cache, about 0.31 MiB per token of context per sequence (~10 GB per sequence at 32k context), plus 10–20 percent overhead.
Is the H100 or H200 better for inference?
For memory-bound serving, the H200: identical compute to the H100 SXM but 141 GB vs. 80 GB and 4.8 vs. 3.35 TB/s bandwidth, which converts directly into throughput at serving batch sizes. As of August 2026 its catalog floor is $2.89/hr vs. $1.99/hr, the cost-per-token math frequently favors it anyway, especially where one H200 replaces two H100s.
What is the cheapest GPU for LLM inference?
For models up to ~14B, marketplace RTX 4090s (from ~$0.59/hr as of August 2026) deliver the lowest cost per token by a wide margin. The cheapest GPU
overall is whichever card your model saturates, compute
price per hour ÷ (tokens per second × 3,600) × 10⁶ for your measured
throughput and compare.
Does hourly GPU price matter at all?
Yes, as the numerator in cost per token, and as the tiebreaker between configurations with similar throughput. It is simply not the ranking metric on its own: the cheapest-per-hour GPU is regularly the most expensive per token once throughput enters the equation.



