Home/GPU Cloud
GPU cloud & AI supercomputingNVIDIA GPU cloud, from one node to 4,096 GPUs
Provision single-tenant NVIDIA accelerators by the second, or reserve an entire InfiniBand-connected cluster for the duration of your training run. Same hardware, same fabric, same engineers — whichever way you buy.
Choose your accelerator
Every SKU is available as a single-tenant bare-metal instance or a virtualised instance with GPU passthrough. Multi-node SKUs ship with rail-optimised InfiniBand as standard.
| GPU | Architecture | Memory / bandwidth | Interconnect | Best for | From (GPU-hr) |
|---|---|---|---|---|---|
| GB300 NVL72 | Blackwell Ultra | 288 GB HBM3e · 8 TB/s | NVLink 5 rack-scale + XDR 800G | Trillion-parameter training, long-context reasoning inference | $8.91 |
| B300 SXM (HGX) | Blackwell Ultra | 288 GB HBM3e · 8 TB/s | NVLink 5 + XDR 800G | Frontier pre-training, MoE fine-tuning | $8.34 |
| B200 SXM (HGX) | Blackwell | 180 GB HBM3e · 8 TB/s | NVLink 5 + NDR/XDR | Large-scale training, FP4 inference | $6.61 |
| H200 SXM (HGX) | Hopper | 141 GB HBM3e · 4.8 TB/s | NVLink 4 + NDR 400G | 70B–400B training and high-throughput serving | $3.79 |
| H100 SXM (HGX) | Hopper | 80 GB HBM3 · 3.35 TB/s | NVLink 4 + NDR 400G | Mainstream training, fine-tuning, inference | $3.10 |
| H100 NVL / PCIe | Hopper | 94 / 80 GB · 3.9 TB/s | NVLink bridge + 200G | Inference density, mixed workloads | $2.70 |
| A100 SXM 80GB | Ampere | 80 GB HBM2e · 2.0 TB/s | NVLink 3 + HDR 200G | Cost-optimised training and batch inference | $1.90 |
| A100 SXM 40GB | Ampere | 40 GB HBM2 · 1.6 TB/s | NVLink 3 + HDR 200G | Cost-optimised training and batch inference | $1.47 |
| RTX PRO 6000 Blackwell | Blackwell | 96 GB GDDR7 · 1.8 TB/s | PCIe Gen5 · 200G | Inference, simulation, graphics & Omniverse | $2.18 |
| L40S | Ada Lovelace | 48 GB GDDR6 · 864 GB/s | PCIe Gen4 · 100G | Inference, rendering, video, VDI | $1.35 |
| RTX 5090 | Blackwell | 32 GB GDDR7 · 1.8 TB/s | PCIe Gen5 · 100G | Research, diffusion, development clusters | $0.75 |
Indicative on-demand list pricing in USD per GPU-hour, exclusive of tax. Reserved and committed terms are discounted; see pricing for term tables.
Buy compute the way your project is funded
On-Demand
Per-second billing, no commitment. Ideal for experimentation, evaluation harnesses and burst fine-tuning.
- 1–8 GPUs per instance
- Instant API / console provisioning
- Snapshot & resume
- No egress charges
Reserved Clusters
Dedicated InfiniBand-connected pods held for 1–36 months, with named capacity and a fixed rate for the term.
- 16–4,096 GPUs, non-blocking fabric
- Managed Slurm or Kubernetes
- Up to 55% below on-demand
- Capacity guarantee in contract
Private AI Cloud
A physically isolated hall, fabric and storage estate operated by Velyrix under your naming, policies and audit regime.
- Dedicated cage or suite
- Customer-managed encryption keys
- Air-gapped option
- Custom SLA & change control
Rail-optimised networking that keeps GPUs busy
A training cluster is only as fast as its slowest all-reduce. Velyrix builds every multi-node pod on a rail-optimised, non-blocking fat-tree with dedicated storage and management planes.
- Compute fabric: NVIDIA Quantum-2 NDR 400G or Quantum-X800 XDR 800G InfiniBand, 1:1 subscription
- Ethernet option: NVIDIA Spectrum-X with RoCEv2, adaptive routing and congestion control
- In-network compute: SHARP aggregation to cut all-reduce latency on large collectives
- Storage plane: separate 200/400G network so checkpoints never contend with gradients
- Management plane: out-of-band BMC network, isolated from tenant traffic
- Validation: NCCL bus bandwidth and all-reduce latency reported before handover
# 512x NVIDIA H200 SXM — XDR 800G rail-optimised nccl-tests/all_reduce_perf -b 8 -e 8G -f 2 -g 8 size busbw algbw status 1.00 GB 372.4 GB/s 198.1 GB/s PASS 4.00 GB 381.9 GB/s 203.7 GB/s PASS 8.00 GB 384.6 GB/s 205.1 GB/s PASS fabric : 1:1 non-blocking, 0 link errors / 72h gpu_burnin : 72h @ 100% — 0 XID, 0 ECC row remaps sustained : 48.9% MFU, Llama-class 70B reference run
Representative figures from a Velyrix acceptance test. Results vary with topology, model, batch size and software stack; your cluster is measured and reported individually.
Arrive with your stack, not a migration project
Managed Slurm
Pre-built Slurm with Pyxis/Enroot, fair-share accounting, topology-aware scheduling and checkpoint-restart hooks for long pre-training runs.
Managed Kubernetes
CNCF-conformant clusters with the NVIDIA GPU Operator, Network Operator, MIG profiles, KubeRay and Kueue for multi-tenant scheduling.
Bare-metal API
Terraform provider and REST API for image deployment, iPXE boot, BMC control and fabric partitioning — build your own control plane on top.
Container registry
Region-local registry and cache for NGC, Docker Hub and private images, so 40 GB pulls do not stall a 512-GPU job launch.
Observability
DCGM, Prometheus and Grafana with per-job GPU utilisation, power, thermals, XID events and fabric counters — exportable to your own SIEM.
Frameworks
Validated images for PyTorch, JAX, NeMo, Megatron-LM, DeepSpeed, TorchTitan, vLLM, SGLang and TensorRT-LLM, refreshed monthly.
What you get before the first job runs
High-performance parallel file system (WEKA or VAST) sized to your token budget, plus S3-compatible object storage for datasets and checkpoints. 10 TB NVMe scratch per node included.
Non-blocking InfiniBand or Spectrum-X compute fabric, dual 100G internet transit, private interconnect to AWS, Azure, GCP and Oracle, and BGP/IP transit options. No egress charges on standard plans.
Single-tenant hosts, isolated VLAN/VRF and fabric partitions, secure-boot firmware attestation, customer-managed keys, and full wipe-and-verify on decommission (NIST SP 800-88 purge).
24×7 NOC, named solutions architect, 15-minute P1 response, hardware replacement targets of four hours on site, and monthly SLA reporting against contracted availability.
Before handover: 72-hour GPU burn-in, memory and ECC validation, NCCL bandwidth tests, fabric error sweep and a signed acceptance report you can hand to your own auditors.
Own the GPUs. Let us run them.
Buy your NVIDIA servers from any OEM or distributor you like, ship them to a Velyrix hall, and we handle the rest — deployment, fabric, cooling, monitoring and support. Or rent ours. Either way, you get a plan in one business day.
Frequently asked questions
How is Velyrix GPU Cloud billed?
On-demand instances bill per second with no minimum term and no data egress charges on standard plans. Reserved clusters bill monthly at a fixed contracted rate for terms from one to thirty-six months. Enterprise private clouds are quoted as a fixed monthly platform fee plus capacity.
Do I get bare metal or virtual machines?
Both are available. The default for multi-node training is single-tenant bare metal with direct access to the GPUs, NICs and NVMe. Virtualised instances with GPU passthrough are offered where snapshotting and rapid re-imaging matter more than the last few percent of performance.
What network do multi-node training clusters use?
Rail-optimised non-blocking InfiniBand - NVIDIA Quantum-2 NDR at 400G or Quantum-X800 XDR at 800G per GPU - with SHARP in-network reduction. NVIDIA Spectrum-X Ethernet with RoCEv2 is available where an Ethernet-only operating model is required.
Can I run Slurm and Kubernetes on the same cluster?
Yes. Velyrix commonly partitions a reserved cluster so that a Slurm partition handles long training runs while a Kubernetes partition serves inference, with a shared parallel file system across both. Partition sizes can be adjusted during the term.
Is there a minimum commitment for large clusters?
Clusters above 256 GPUs are normally contracted for a minimum of three months because of fabric build and cabling work. Shorter windows can be accommodated in halls where a matching pod is already built and idle.