New Now onboarding NVIDIA GB300 NVL72 clusters
GPU infrastructure for ML teams

High-performance compute,
cloud simplicity

Run and scale AI workloads on bare-metal-level GPU instances - from a single GPU to thousand-GPU clusters - or serve models on production inference endpoints. One platform, predictable pricing, no cluster wrangling.

$2.99 per GPU-hr, HGX H100 on-demand
3.2 Tb/s InfiniBand per GPU node
99.9% platform uptime target

Runs with the tools your team already uses

PyTorch vLLM Hugging Face TensorFlow JAX Ray NVIDIA NGC
The platform

Two ways to run, one platform

Rent raw GPU power by the hour, or hand us a model and get an endpoint. Same fabric, same billing, no cluster to babysit.

Compute

GPU instances and clusters for training, fine-tuning, and batch jobs - from a single card to thousand-GPU runs.

  • Bare-metal-level instances, no GPU or network virtualization
  • Non-blocking InfiniBand fabric across every node
  • Billed per GPU-hour, on-demand or reserved
Explore compute

Inference

Deploy open or custom models behind a single OpenAI-compatible endpoint. We handle GPU allocation, autoscaling, and health checks.

  • Drop-in OpenAI-compatible API - change one base URL
  • Autoscaling with scale-to-zero on idle endpoints
  • Pay per token, no reserved-capacity contracts
Explore inference
$ curl https://api.meridian.finimble.com/v1/chat/completions \
  -H "Authorization: Bearer $MERIDIAN_API_KEY" \
  -d '{
    "model": "llama-3.1-70b",
    "messages": [{"role":"user","content":"Hello"}]
  }'
Why Meridian

Supercomputer performance, without the ops team.

Built for scale

Engineered so your GPUs actually run your model

Every layer - silicon, fabric, and scheduling - is tuned to keep expensive accelerators busy and your runs alive.

Bare-metal-level performance

No GPU or network virtualization means up to 20% higher system MFU than comparative benchmarks - less infrastructure for the same result.

  • Direct access to SXM GPUs and NICs
  • Quantum-2 InfiniBand, no oversubscription

Resilient from the start

Automated health checks and node lifecycle management mean 50% fewer interruptions per day across your fleet - long runs stay alive.

  • Continuous fleet health monitoring
  • Automatic drain, replace, and reschedule

Developer-first

An OpenAI-compatible API, a real CLI, and docs that get you from signup to first request in minutes - not a support ticket.

  • Provision instances via CLI or API
  • Bring your own images and checkpoints
GPU lineup

Accelerated compute, powered by NVIDIA

Rack-scale GB300 NVL72 for frontier training, Blackwell and Hopper HGX nodes for large-scale workloads, and CPU instances for the rest of your pipeline.

  • GB300 NVL72
  • HGX B200
  • HGX H200
  • HGX H100
  • CPU instances
Explore the full lineup
By the numbers

Infrastructure that earns its keep

20% higher system MFU vs. comparative benchmarks
50% fewer interruptions per day across the fleet
3.2 Tb/s InfiniBand bandwidth per GPU node

Launch your first GPU in minutes

Start from the quickstart, or talk to our team about capacity and reserved pricing.