High-performance compute,
cloud simplicity
Run and scale AI workloads on bare-metal-level GPU instances - from a single GPU to thousand-GPU clusters - or serve models on production inference endpoints. One platform, predictable pricing, no cluster wrangling.
Runs with the tools your team already uses
Two ways to run, one platform
Rent raw GPU power by the hour, or hand us a model and get an endpoint. Same fabric, same billing, no cluster to babysit.
Compute
GPU instances and clusters for training, fine-tuning, and batch jobs - from a single card to thousand-GPU runs.
- Bare-metal-level instances, no GPU or network virtualization
- Non-blocking InfiniBand fabric across every node
- Billed per GPU-hour, on-demand or reserved
Inference
Deploy open or custom models behind a single OpenAI-compatible endpoint. We handle GPU allocation, autoscaling, and health checks.
- Drop-in OpenAI-compatible API - change one base URL
- Autoscaling with scale-to-zero on idle endpoints
- Pay per token, no reserved-capacity contracts
$ curl https://api.meridian.finimble.com/v1/chat/completions \
-H "Authorization: Bearer $MERIDIAN_API_KEY" \
-d '{
"model": "llama-3.1-70b",
"messages": [{"role":"user","content":"Hello"}]
}'
Supercomputer performance, without the ops team.
Engineered so your GPUs actually run your model
Every layer - silicon, fabric, and scheduling - is tuned to keep expensive accelerators busy and your runs alive.
Bare-metal-level performance
No GPU or network virtualization means up to 20% higher system MFU than comparative benchmarks - less infrastructure for the same result.
- Direct access to SXM GPUs and NICs
- Quantum-2 InfiniBand, no oversubscription
Resilient from the start
Automated health checks and node lifecycle management mean 50% fewer interruptions per day across your fleet - long runs stay alive.
- Continuous fleet health monitoring
- Automatic drain, replace, and reschedule
Developer-first
An OpenAI-compatible API, a real CLI, and docs that get you from signup to first request in minutes - not a support ticket.
- Provision instances via CLI or API
- Bring your own images and checkpoints
Accelerated compute, powered by NVIDIA
Rack-scale GB300 NVL72 for frontier training, Blackwell and Hopper HGX nodes for large-scale workloads, and CPU instances for the rest of your pipeline.
- GB300 NVL72
- HGX B200
- HGX H200
- HGX H100
- CPU instances
Infrastructure that earns its keep
Launch your first GPU in minutes
Start from the quickstart, or talk to our team about capacity and reserved pricing.