Pricing
Pay for what you use. Nothing else.
Per-GPU-hour compute, per-token inference, per-vCPU CPU instances. No egress fees, no minimums, no reserved-capacity lock-in.
| GPU | Memory | On-demand | Reserved (1-yr) | Best for |
|---|---|---|---|---|
| GB300 NVL72 | Rack-scale, liquid-cooled | Contact sales | Contact sales | Frontier training, AI reasoning |
| HGX B200 | 180 GB HBM3e / GPU | $5.50 / GPU-hr | $3.85 / GPU-hr | Large-scale LLM training, MoE |
| HGX H200 | 141 GB HBM3e / GPU | $3.99 / GPU-hr | $2.79 / GPU-hr | Memory-intensive inference |
| HGX H100 | 80 GB HBM3 / GPU | $2.99 / GPU-hr | $2.09 / GPU-hr | Cost-efficient training & inference |
All GPU instances include local NVMe scratch storage. Multi-GPU nodes include InfiniBand interconnect at no extra charge.
| Model class | Input | Output | Notes |
|---|---|---|---|
| Small (7-13B) | $0.10 / 1M tokens | $0.20 / 1M tokens | Scale-to-zero eligible |
| Medium (30-40B) | $0.40 / 1M tokens | $0.80 / 1M tokens | Scale-to-zero eligible |
| Large (70B+) | $0.90 / 1M tokens | $1.80 / 1M tokens | Dedicated capacity available |
| Instance | Configuration | Price | Best for |
|---|---|---|---|
| cpu-standard | Intel Xeon or AMD EPYC | $0.048 / vCPU-hr | App backends, agents, orchestration |
| cpu-highmem | 8 GB RAM / vCPU | $0.064 / vCPU-hr | Data preprocessing, feature engineering |
FAQ
Pricing questions, answered
How granular is billing?
GPU and CPU instances are metered per second and billed per hour of usage. Stop the instance, stop the bill. Inference is billed per token.
Are there egress or storage fees?
No egress fees. Local NVMe scratch is included with every GPU instance. Persistent block storage is billed separately per GB-month.
What does reserved capacity commit me to?
Reserved pricing reflects a 1-year commitment on a fixed number of GPUs, in exchange for roughly 30% off on-demand rates. Capacity is guaranteed for the term.
Do you offer preemptible / spot capacity?
Yes - preemptible instances are available on the H100 tier at a significant discount, for fault-tolerant batch workloads.
Can I bring my own model or checkpoint?
Yes. Inference endpoints can serve your own fine-tuned checkpoints alongside the open-weight model catalog, at the same per-token rates.
Need committed capacity or custom terms?
Volume discounts, reserved GPU pools, and private deployments are available for teams with predictable workloads.