Bring your own GPUs. We build the AI factory around them.
Purchase your NVIDIA GPU servers from any authorized OEM or distributor, ship them directly to a Velyrix facility, and our engineering team handles deployment, networking, cooling, cluster configuration, monitoring and ongoing support. Prefer not to own hardware? Rent the same class of cluster from our fleet instead. Either way, one partner from the loading dock to the running model.
Bring Your Own GPU: you buy the hardware, we run it for as long as you own it
Purchase your NVIDIA GPU servers from any authorized OEM or distributor — Dell, Supermicro, GIGABYTE, HPE, Lenovo or anyone else. Ship them directly to our facility, and our engineering team will handle deployment, networking, cooling, cluster configuration, monitoring, and ongoing support. We do not make money on your servers. We make money keeping them running.
Optional procurement assistance is available — configuration review, vendor-neutral platform comparison and lead-time intelligence — but you are never required to buy hardware through Velyrix, and your pricing is identical either way.
Nine ways to work with Velyrix
Bring your own hardware into our halls, rent by the hour, lease a private cluster, have us build and commission the whole AI factory — or sell all of it under your own brand.
Bring Your Own GPU
You buy the NVIDIA servers from any OEM and keep the asset. We deploy, cool, network, monitor and support them — with no hardware margin and no lock-in.
- Any authorized OEM or distributor
- Warranty-preserving installation
- Optional procurement assistance
GPU Cloud
On-demand and reserved NVIDIA GPU instances with InfiniBand-connected multi-node training, Slurm or Kubernetes, and per-second billing.
- H100, H200, B200, B300 & GB300
- 8 – 4,096 GPU clusters
- Spin-up in under 90 seconds
Dedicated GPU Servers
Single-tenant bare metal by the month or year. Root access, no hypervisor tax, no noisy neighbours, transparent published pricing.
- 8×H100 / H200 / B200 nodes
- A100, L40S, RTX PRO 6000
- 1 – 36 month terms
AI Colocation
Bring your own HGX, DGX, MGX or OEM GPU systems into liquid-cooled halls engineered for 40–160 kW racks.
- Air, RDHx and direct-to-chip liquid
- Cage, private suite or wholesale
- Remote hands 24×7
AI Infrastructure Deployment & Commissioning
Turnkey build-out: site survey, rack & stack, fabric cabling, firmware, burn-in, NCCL validation and documented acceptance testing.
- Greenfield AI factory delivery
- InfiniBand & Spectrum-X fabric build
- Signed acceptance test reports
Private LLM Deployment & Optimization
Your own DeepSeek, Qwen, Llama, GPT-OSS, Mistral, GLM, Kimi, MiniMax, Gemma or Nemotron endpoint — tuned, quantised and air-gappable.
- vLLM, SGLang & TensorRT-LLM serving
- FP8 / FP4 / AWQ quantisation
- LoRA fine-tuning & RAG
Storage, Network & Managed Ops
Parallel file systems, object storage, 400G transit, private interconnect and a 24×7 NOC that runs the cluster for you.
- WEKA / VAST / Ceph tiers
- Managed Slurm & Kubernetes
- Observability and SLA reporting
Migration, Audit & Lifecycle
Move a live GPU estate between halls or providers, or have an independent engineering team tell you why your existing cluster underperforms.
- Staged migration with minimal downtime
- Fabric, firmware & MFU health audit
- Refresh and certified decommission
OEM & Channel Partners
We are the deployment and support bench behind hardware you sell — under your brand, with deal registration and a written channel-neutrality policy.
- White-label deployment & ATP
- L1–L3 support, RMA and spares
- We never resell servers ourselves
Every generation of NVIDIA accelerator, in one operator
From Ampere for cost-optimised inference to Blackwell Ultra racks for trillion-parameter training, Velyrix runs a single operational model across the entire fleet — the same fabric design, the same monitoring, the same SLA.
- Blackwell Ultra: NVIDIA GB300 NVL72 and HGX B300 in direct-to-chip liquid-cooled halls
- Blackwell: HGX B200 8-GPU nodes with 1.4 TB HBM3e per node
- Hopper: HGX H200 and H100 SXM at scale, available today in every region
- Ampere & Ada: A100 80GB, L40S and RTX PRO 6000 Blackwell for inference and visualisation
Indicative rack and node power envelopes used for capacity planning. Actual draw depends on OEM chassis, cooling method and workload profile.
Transparent GPU pricing, published up front
Starting rates for single-tenant capacity. Multi-node reserved clusters and 12–36 month terms are quoted individually — larger commitments move well below these numbers.
| Accelerator | Memory | On-demand | 6-month reserved | 8-GPU node / month | Status |
|---|---|---|---|---|---|
| GB300 NVL72 | 288 GB HBM3e | from $8.91 | $8.20 | Rack-scale — quoted | Q4 2026 |
| B300 SXM | 288 GB HBM3e | from $8.34 | $7.67 | from $44,800 | Limited |
| B200 SXM | 180 GB HBM3e | from $6.61 | $6.08 | from $38,000 | Available |
| H200 SXM | 141 GB HBM3e | from $3.79 | $3.49 | from $23,900 | Available |
| H100 SXM | 80 GB HBM3 | from $3.10 | $2.85 | from $18,400 | Available |
| A100 SXM 80GB | 80 GB HBM2e | from $1.90 | $1.75 | from $10,400 | Available |
Indicative list pricing in USD, exclusive of tax, effective September 2026. Prices vary by region, term, storage and egress profile, and are confirmed in a written quotation. See the full price list.
Run open-weight frontier models inside your own perimeter
Velyrix deploys, quantises and optimises the leading open-weight model families on dedicated Velyrix hardware, in your colocation footprint, or fully air-gapped on-premises — with an OpenAI-compatible API and measured tokens-per-second acceptance criteria.
# Your private endpoint — no data leaves your VPC curl https://llm.internal.acme.com/v1/chat/completions \ -H "Authorization: Bearer $VELYRIX_KEY" \ -d '{ "model": "deepseek-v3-fp8", "messages": [{"role":"user","content":"Summarise Q3 risk report"}], "max_tokens": 1024 }' # Served by vLLM on 8x NVIDIA H200 — measured at handover throughput : 14,850 tok/s aggregate ttft_p95 : 218 ms availability: 99.9% monthly SLA
25 markets across 18 countries — every one of them export-clean
Velyrix places capacity only in jurisdictions where advanced NVIDIA accelerators can be deployed lawfully and without licensing risk, inside carrier-neutral, independently audited partner data centres. Two markets are live with capacity held today; the rest are delivered on request against your specific requirement. We tell you which is which before you commit — see the full status table.
"Live" means Velyrix holds contracted space and power today. "On request" means we contract space against your specific requirement, adding roughly two to six weeks. "Planned" means still in negotiation — we will not take an order against it. Every market on this list is a permitted destination for the accelerators we deploy, and every order is screened before hardware ships; see trade compliance and how we choose countries.
Whatever you buy, we know how to deploy it
Velyrix holds validated deployment playbooks — rack elevation, power schedule, cooling method, cable plan and acceptance profile — for the platforms the market actually ships.
Also NVIDIA DGX, ASUS, QCT, Wiwynn and Foxconn ORv3 designs. Platform names are trademarks of their respective owners, listed to describe deployment capability. Are you an OEM, distributor or reseller? See our partner programme.
Engineered like a hyperscaler, responsive like a specialist
Weeks, not quarters
Reserved clusters delivered in 4–8 weeks; on-demand capacity live in minutes. Greenfield AI factories commissioned in 12–20 weeks.
Goodput, not just FLOPS
We contract on measured cluster goodput. Every handover includes NCCL bus-bandwidth and sustained MFU acceptance numbers — and if we miss the agreed figure we fix it at our cost.
Single tenant by default
Dedicated hosts, dedicated fabric partitions, customer-held encryption keys and optional air-gapped enclaves.
Efficient by design
Design PUE of 1.15–1.25 with direct-to-chip liquid cooling, heat reuse where available and renewable power contracting.
We would rather fill this space with commitments than with logos. So here is what goes into your contract: measured acceptance criteria tested before you pay, service credits that apply automatically without a claim process, month-to-month terms available on BYOG hosting, and the direct phone number of the engineer who built your cluster.
Frequently asked questions
Can I bring GPU servers I bought myself?
Yes - that is our flagship Bring Your Own GPU model. Purchase NVIDIA GPU servers from any authorized OEM or distributor, ship them directly to a Velyrix facility, and our engineering team handles deployment, networking, cooling, cluster configuration, monitoring and ongoing support. You keep ownership, the warranty and the OEM relationship; optional procurement assistance is available but never required.
What is a neocloud, and how is Velyrix different from a hyperscaler?
A neocloud is a cloud provider purpose-built for accelerated computing rather than general-purpose IT. Velyrix operates only AI infrastructure: high-density liquid-cooled halls, NVIDIA GPU fleets, lossless InfiniBand and Spectrum-X fabrics and parallel storage. That focus means higher GPU density per rack, non-blocking fabrics as standard, published pricing and direct access to the engineers who built your cluster.
Which NVIDIA GPUs can I rent or host with Velyrix?
Velyrix supports NVIDIA GB300 NVL72, HGX B300, HGX B200, HGX H200, HGX H100, HGX A100, L40S and RTX PRO 6000 Blackwell systems from the major OEMs, available as GPU cloud instances, dedicated bare-metal servers or customer-owned hardware hosted in Velyrix colocation.
Can I bring my own GPU servers instead of renting?
Yes. Velyrix colocation is engineered for 40-160 kW racks with air, rear-door heat exchanger and direct-to-chip liquid cooling, and accepts HGX, DGX, MGX and OCP ORv3 systems from Dell, Supermicro, HPE, Lenovo, Gigabyte, QCT and others. We can also deploy and commission the hardware for you.
How quickly can a cluster be delivered?
On-demand GPU cloud instances start in under 90 seconds. Reserved multi-node clusters in an existing hall are typically live in 4-8 weeks. A greenfield AI factory build, including fabric, cooling and commissioning, is normally 12-20 weeks depending on hardware lead times.
Do you help with deploying our own LLM rather than using a public API?
Yes. Our Private LLM Deployment and Optimization service installs, quantises, benchmarks and operates open-weight models such as DeepSeek, Qwen, Llama, OpenAI gpt-oss, Mistral, GLM, Kimi, MiniMax, Gemma and NVIDIA Nemotron on infrastructure you control, exposed through an OpenAI-compatible API with contracted throughput and latency targets.