New Now onboarding NVIDIA GB300 NVL72 clusters

Pricing

Pay for what you use. Nothing else.

Per-GPU-hour compute, per-token inference, per-vCPU CPU instances. No egress fees, no minimums, no reserved-capacity lock-in.

GPU compute pricing
GPU Memory On-demand Reserved (1-yr) Best for
GB300 NVL72 Rack-scale, liquid-cooled Contact sales Contact sales Frontier training, AI reasoning
HGX B200 180 GB HBM3e / GPU $5.50 / GPU-hr $3.85 / GPU-hr Large-scale LLM training, MoE
HGX H200 141 GB HBM3e / GPU $3.99 / GPU-hr $2.79 / GPU-hr Memory-intensive inference
HGX H100 80 GB HBM3 / GPU $2.99 / GPU-hr $2.09 / GPU-hr Cost-efficient training & inference

All GPU instances include local NVMe scratch storage. Multi-GPU nodes include InfiniBand interconnect at no extra charge.

Inference pricing
Model class Input Output Notes
Small (7-13B) $0.10 / 1M tokens $0.20 / 1M tokens Scale-to-zero eligible
Medium (30-40B) $0.40 / 1M tokens $0.80 / 1M tokens Scale-to-zero eligible
Large (70B+) $0.90 / 1M tokens $1.80 / 1M tokens Dedicated capacity available
CPU instance pricing
Instance Configuration Price Best for
cpu-standard Intel Xeon or AMD EPYC $0.048 / vCPU-hr App backends, agents, orchestration
cpu-highmem 8 GB RAM / vCPU $0.064 / vCPU-hr Data preprocessing, feature engineering

FAQ

Pricing questions, answered

How granular is billing?

GPU and CPU instances are metered per second and billed per hour of usage. Stop the instance, stop the bill. Inference is billed per token.

Are there egress or storage fees?

No egress fees. Local NVMe scratch is included with every GPU instance. Persistent block storage is billed separately per GB-month.

What does reserved capacity commit me to?

Reserved pricing reflects a 1-year commitment on a fixed number of GPUs, in exchange for roughly 30% off on-demand rates. Capacity is guaranteed for the term.

Do you offer preemptible / spot capacity?

Yes - preemptible instances are available on the H100 tier at a significant discount, for fault-tolerant batch workloads.

Can I bring my own model or checkpoint?

Yes. Inference endpoints can serve your own fine-tuned checkpoints alongside the open-weight model catalog, at the same per-token rates.

Need committed capacity or custom terms?

Volume discounts, reserved GPU pools, and private deployments are available for teams with predictable workloads.