Home/Dedicated Servers

Bare metal & dedicated GPU servers

Single-tenant GPU servers, published prices, no hypervisor tax

Lease a whole machine — every GPU, every core, every NVMe lane, root and BMC access included. Monthly, annual or hourly, across 25 markets, with hardware from Dell, Supermicro, HPE, Lenovo, GIGABYTE and QCT.

Would you rather own the hardware? Buy from any OEM and let us run it — see Bring Your Own GPU. Renting and owning use the same halls, the same fabric and the same engineers.

Tenancy
100% single tenant
Access
Root + IPMI/Redfish BMC
Terms
Hourly · 1–36 months
Provisioning
2 hours – 10 days
Bandwidth
10–100G unmetered options
Flagship AI nodes

NVIDIA HGX & NVL72 dedicated servers

Eight-GPU SXM nodes with NVLink and NVSwitch, dual-socket CPUs, 400G/800G fabric NICs and local NVMe scratch. Sold as whole nodes and as multi-node pods with InfiniBand between them.

ConfigurationCPU / RAMLocal NVMeFabric NICMonthlyAnnual (per mo)Hourly
GB300 NVL72 — 72 GPU36× Grace · 40 TB4× 7.68 TB18× XDR 800GQuotedQuoted
8× B300 SXM 288GB2× Xeon 6 / EPYC 9005 · 3 TB8× 7.68 TB Gen58× XDR 800G$44,800$39,900$66.72
8× B200 SXM 180GB2× Xeon 6 / EPYC 9005 · 3 TB8× 7.68 TB Gen58× NDR/XDR$38,000$33,800$52.88
8× H200 SXM 141GB2× Xeon 8568Y+ / EPYC 9554 · 2 TB8× 3.84 TB Gen4/58× NDR 400G$23,900$21,300$30.32
8× H100 SXM 80GB2× Xeon 8468 / EPYC 9474F · 2 TB8× 3.84 TB Gen48× NDR 400G$18,400$16,400$24.80
8× A100 SXM 80GB2× EPYC 7763 / Xeon 8358 · 1 TB4× 3.84 TB Gen48× HDR 200G$10,400$9,300$15.20
8× A100 SXM 40GB2× EPYC 7742 · 1 TB4× 1.92 TB Gen48× HDR 200G$8,000$7,100$11.76

Indicative list pricing in USD per node, excluding tax, parallel storage and cross-region transit. Multi-node pods include InfiniBand leaf/spine, cabling and validation.

PCIe & inference servers

The configurations the market actually rents

Inference, fine-tuning, rendering and CI fleets rarely need an NVL72. These are the workhorse builds, with the same single-tenant guarantee.

ConfigurationGPU memoryPlatformTypical workloadMonthlyHourly
8× H100 NVL 94GB752 GB2× EPYC 9454 · 1.5 TB · PCIe Gen5High-density LLM inference$14,700$21.60
4× H100 PCIe 80GB320 GB2× Xeon 8462Y+ · 768 GBFine-tuning, mid-size serving$7,300$10.80
8× RTX PRO 6000 Blackwell 96GB768 GB2× EPYC 9455 · 1.5 TB · Gen5Inference, Omniverse, simulation$9,500$17.44
8× L40S 48GB384 GB2× Xeon 8462Y+ · 1 TBInference, video, VDI, rendering$6,000$10.80
4× L40S 48GB192 GB2× EPYC 9354 · 512 GBSmall-model serving, batch$3,300$5.40
8× RTX 5090 32GB256 GB2× EPYC 9354 · 768 GBResearch, diffusion, dev clusters$3,800$6.00
8× RTX 4090 24GB192 GB2× EPYC 7543 · 512 GBDiffusion, CI, cost-sensitive inference$2,700$5.04
4× RTX A6000 48GB192 GB2× Xeon 6338 · 512 GBWorkstation-class, CAD, legacy CUDA$1,900$5.40
8× L4 24GB192 GB2× EPYC 9124 · 384 GBLow-power inference, transcoding$2,200$3.12
2× L40S 48GB96 GB1× EPYC 9254 · 256 GBEntry inference node$1,640$2.70
CPU & storage bare metal

The rest of the estate

Data pipelines, tokenisation, head nodes, vector databases and checkpoint stores — all on the same network, in the same hall as your GPUs.

ConfigurationCores / threadsMemoryStorageNetworkMonthly
Compute — Dual EPYC 9654192C / 384T768 GB DDR52× 3.84 TB NVMe2× 25G$1,490
Compute — Dual Xeon 6740E192C / 192T512 GB DDR52× 1.92 TB NVMe2× 25G$1,210
General — Single EPYC 935432C / 64T256 GB DDR52× 1.92 TB NVMe2× 10G$680
Head / login node16C / 32T128 GB2× 960 GB NVMe2× 10G$370
Storage — NVMe dense64C / 128T512 GB24× 7.68 TB NVMe (184 TB)2× 100G$3,800
Storage — capacity HDD32C / 64T256 GB60× 22 TB SAS (1.3 PB raw)2× 25G$4,500
Memory-optimised96C / 192T3 TB DDR54× 3.84 TB NVMe2× 100G$2,800
Market context

Priced above the median, deliberately

Velyrix list pricing is set at 15% above the market median — deliberately, and we publish the arithmetic. We are not trying to be the cheapest GPU-hour on the internet, because the cheapest GPU-hour is usually oversubscribed, blocking, unsupported, or all three. The premium buys engineering: non-blocking fabric, a signed acceptance report, four-hour hardware replacement and engineers who answer at 3 a.m.

Rental unitMarket rangeMarket medianVelyrix listvs medianWhat the premium buys
H100 SXM — per GPU-hour$1.90 – $3.50$2.70$3.10+15%1:1 non-blocking NDR fabric, single tenancy
H200 SXM — per GPU-hour$2.40 – $4.20$3.30$3.79+15%Guaranteed capacity, no oversubscription
B200 — per GPU-hour$4.00 – $7.50$5.75$6.61+15%Liquid-cooled halls, sustained clocks under load
B300 / GB300 — per GPU-hour$5.50 – $9.00$7.25$8.34+15%Rack-scale commissioning and CDU redundancy
A100 80GB — per GPU-hour$1.10 – $2.20$1.65$1.90+15%Maintained fleet, burn-in before every handover
L40S — per GPU-hour$0.75 – $1.60$1.18$1.35+15%Enterprise SLA rather than best effort
RTX 4090 — per GPU-hour$0.30 – $0.80$0.55$0.63+15%Tier III-class facility, not a colocated shelf
8× H100 SXM 80GB node — per month$10,000 – $22,000$16,000$18,400+15%24×7 NOC, 4-hour replacement, acceptance report
8× H200 SXM 141GB node — per month$13,500 – $28,000$20,750$23,900+15%Named architect, monthly SLA reporting
8× B200 SXM 180GB node — per month$21,000 – $45,000$33,000$38,000+15%Direct-to-chip cooling, validated fabric

Market ranges and medians are Velyrix's own observation of publicly advertised list pricing across GPU cloud and bare-metal providers, provided for orientation only; they are not a representation about any third party's current pricing, and no comparison is endorsed by any other provider. Velyrix rates are indicative list prices confirmed in a written quotation. Committed multi-year terms and large volumes are quoted individually.

Included & optional

What ships with every dedicated server

  • Root and BMC access: IPMI/Redfish, virtual media, serial-over-LAN and remote reboot
  • OS of your choice: Ubuntu, Rocky, RHEL, Debian, Windows Server or a custom image via iPXE
  • Driver stack pre-validated: NVIDIA driver, CUDA, Fabric Manager, DCGM, NVIDIA Container Toolkit
  • Burn-in before handover: 24–72 hour GPU stress, ECC and thermal validation report
  • Bandwidth: 10G unmetered or 25/100G metered, dual-homed, with DDoS mitigation
  • Remote hands: 24×7 facility engineers on site, four-hour hardware replacement target
Add-onUnitPrice
Parallel FS (WEKA/VAST)per TB / month$42
NVMe block storageper TB / month$100
Object storage (S3 API)per TB / month$20
Backup & snapshot vaultper TB / month$14
100G unmetered uplinkper port / month$1,600
Additional IPv4 /29per month$46
Private cloud interconnectper 10G / month$545
InfiniBand pod networkingper node / month$805
Managed OS & patchingper server / month$230
Managed Slurm / Kubernetesper cluster / month$2,400
Talk to an AI infrastructure architect

Own the GPUs. Let us run them.

Buy your NVIDIA servers from any OEM or distributor you like, ship them to a Velyrix hall, and we handle the rest — deployment, fabric, cooling, monitoring and support. Or rent ours. Either way, you get a plan in one business day.

Frequently asked questions

What is the difference between a dedicated GPU server and a GPU cloud instance?

A dedicated server is a whole physical machine leased to one customer: all GPUs, all CPU cores, all NVMe and the BMC belong to you for the term, with no hypervisor overhead and no shared tenancy. A cloud instance is provisioned per second and may be virtualised. Dedicated servers are usually 25-45 percent cheaper per GPU-hour on terms of three months or more.

How much does it cost to rent an 8x H100 server per month?

Velyrix lists an 8x NVIDIA H100 SXM node from $18,400 per month, falling to $16,400 on a twelve-month term and $15,600 at thirty-six months. The list price is deliberately 15 percent above the market median of roughly $16,000: it reflects non-blocking fabric, a signed acceptance report, 24x7 NOC coverage and a four-hour hardware replacement target rather than best-effort hosting.

Can I rent a single GPU rather than a whole node?

Yes, through Velyrix GPU Cloud, which offers 1, 2, 4 and 8 GPU instances billed per second. Dedicated servers are sold as whole nodes because the NVLink and NVSwitch topology inside an HGX baseboard cannot be meaningfully divided between tenants.

How quickly is a dedicated server delivered?

Configurations held in stock in an active region are typically live within two to twenty-four hours. Custom builds, large pods and Blackwell-class nodes depend on OEM lead times and are normally delivered within five to twenty business days.

Do you charge for bandwidth or egress?

Standard plans include a 10G unmetered uplink with no egress charges. Higher-capacity 25G and 100G ports, private cloud interconnects and dedicated transit are priced as add-ons.

Can I upgrade or swap hardware mid-term?

Yes. Within a term you can move to a higher tier at the difference in rate, keeping your data on the same storage estate. Downgrades are applied at the next renewal date.