Home/GPU Cloud

GPU cloud & AI supercomputing

NVIDIA GPU cloud, from one node to 4,096 GPUs

Provision single-tenant NVIDIA accelerators by the second, or reserve an entire InfiniBand-connected cluster for the duration of your training run. Same hardware, same fabric, same engineers — whichever way you buy.

Provisioning
< 90 seconds
Billing
Per second, no egress fee
Fabric
400G / 800G non-blocking
Max cluster
4,096 GPUs single fabric
Uptime SLA
99.99% infrastructure
GPU catalogue

Choose your accelerator

Every SKU is available as a single-tenant bare-metal instance or a virtualised instance with GPU passthrough. Multi-node SKUs ship with rail-optimised InfiniBand as standard.

GPUArchitectureMemory / bandwidthInterconnectBest forFrom (GPU-hr)
GB300 NVL72Blackwell Ultra288 GB HBM3e · 8 TB/sNVLink 5 rack-scale + XDR 800GTrillion-parameter training, long-context reasoning inference$8.91
B300 SXM (HGX)Blackwell Ultra288 GB HBM3e · 8 TB/sNVLink 5 + XDR 800GFrontier pre-training, MoE fine-tuning$8.34
B200 SXM (HGX)Blackwell180 GB HBM3e · 8 TB/sNVLink 5 + NDR/XDRLarge-scale training, FP4 inference$6.61
H200 SXM (HGX)Hopper141 GB HBM3e · 4.8 TB/sNVLink 4 + NDR 400G70B–400B training and high-throughput serving$3.79
H100 SXM (HGX)Hopper80 GB HBM3 · 3.35 TB/sNVLink 4 + NDR 400GMainstream training, fine-tuning, inference$3.10
H100 NVL / PCIeHopper94 / 80 GB · 3.9 TB/sNVLink bridge + 200GInference density, mixed workloads$2.70
A100 SXM 80GBAmpere80 GB HBM2e · 2.0 TB/sNVLink 3 + HDR 200GCost-optimised training and batch inference$1.90
A100 SXM 40GBAmpere40 GB HBM2 · 1.6 TB/sNVLink 3 + HDR 200GCost-optimised training and batch inference$1.47
RTX PRO 6000 BlackwellBlackwell96 GB GDDR7 · 1.8 TB/sPCIe Gen5 · 200GInference, simulation, graphics & Omniverse$2.18
L40SAda Lovelace48 GB GDDR6 · 864 GB/sPCIe Gen4 · 100GInference, rendering, video, VDI$1.35
RTX 5090Blackwell32 GB GDDR7 · 1.8 TB/sPCIe Gen5 · 100GResearch, diffusion, development clusters$0.75

Indicative on-demand list pricing in USD per GPU-hour, exclusive of tax. Reserved and committed terms are discounted; see pricing for term tables.

Consumption models

Buy compute the way your project is funded

Elastic

On-Demand

Per-second billing, no commitment. Ideal for experimentation, evaluation harnesses and burst fine-tuning.

  • 1–8 GPUs per instance
  • Instant API / console provisioning
  • Snapshot & resume
  • No egress charges
Most popular

Reserved Clusters

Dedicated InfiniBand-connected pods held for 1–36 months, with named capacity and a fixed rate for the term.

  • 16–4,096 GPUs, non-blocking fabric
  • Managed Slurm or Kubernetes
  • Up to 55% below on-demand
  • Capacity guarantee in contract
Enterprise

Private AI Cloud

A physically isolated hall, fabric and storage estate operated by Velyrix under your naming, policies and audit regime.

  • Dedicated cage or suite
  • Customer-managed encryption keys
  • Air-gapped option
  • Custom SLA & change control
Cluster fabric

Rail-optimised networking that keeps GPUs busy

A training cluster is only as fast as its slowest all-reduce. Velyrix builds every multi-node pod on a rail-optimised, non-blocking fat-tree with dedicated storage and management planes.

  • Compute fabric: NVIDIA Quantum-2 NDR 400G or Quantum-X800 XDR 800G InfiniBand, 1:1 subscription
  • Ethernet option: NVIDIA Spectrum-X with RoCEv2, adaptive routing and congestion control
  • In-network compute: SHARP aggregation to cut all-reduce latency on large collectives
  • Storage plane: separate 200/400G network so checkpoints never contend with gradients
  • Management plane: out-of-band BMC network, isolated from tenant traffic
  • Validation: NCCL bus bandwidth and all-reduce latency reported before handover
acceptance-report.txt
# 512x NVIDIA H200 SXM — XDR 800G rail-optimised
nccl-tests/all_reduce_perf -b 8 -e 8G -f 2 -g 8

size        busbw       algbw      status
1.00 GB     372.4 GB/s  198.1 GB/s  PASS
4.00 GB     381.9 GB/s  203.7 GB/s  PASS
8.00 GB     384.6 GB/s  205.1 GB/s  PASS

fabric      : 1:1 non-blocking, 0 link errors / 72h
gpu_burnin  : 72h @ 100% — 0 XID, 0 ECC row remaps
sustained   : 48.9% MFU, Llama-class 70B reference run

Representative figures from a Velyrix acceptance test. Results vary with topology, model, batch size and software stack; your cluster is measured and reported individually.

Orchestration & software

Arrive with your stack, not a migration project

Managed Slurm

Pre-built Slurm with Pyxis/Enroot, fair-share accounting, topology-aware scheduling and checkpoint-restart hooks for long pre-training runs.

Managed Kubernetes

CNCF-conformant clusters with the NVIDIA GPU Operator, Network Operator, MIG profiles, KubeRay and Kueue for multi-tenant scheduling.

Bare-metal API

Terraform provider and REST API for image deployment, iPXE boot, BMC control and fabric partitioning — build your own control plane on top.

📦

Container registry

Region-local registry and cache for NGC, Docker Hub and private images, so 40 GB pulls do not stall a 512-GPU job launch.

📈

Observability

DCGM, Prometheus and Grafana with per-job GPU utilisation, power, thermals, XID events and fabric counters — exportable to your own SIEM.

🔧

Frameworks

Validated images for PyTorch, JAX, NeMo, Megatron-LM, DeepSpeed, TorchTitan, vLLM, SGLang and TensorRT-LLM, refreshed monthly.

Included with every cluster

What you get before the first job runs

Storage

High-performance parallel file system (WEKA or VAST) sized to your token budget, plus S3-compatible object storage for datasets and checkpoints. 10 TB NVMe scratch per node included.

Networking

Non-blocking InfiniBand or Spectrum-X compute fabric, dual 100G internet transit, private interconnect to AWS, Azure, GCP and Oracle, and BGP/IP transit options. No egress charges on standard plans.

Security

Single-tenant hosts, isolated VLAN/VRF and fabric partitions, secure-boot firmware attestation, customer-managed keys, and full wipe-and-verify on decommission (NIST SP 800-88 purge).

Support

24×7 NOC, named solutions architect, 15-minute P1 response, hardware replacement targets of four hours on site, and monthly SLA reporting against contracted availability.

Acceptance

Before handover: 72-hour GPU burn-in, memory and ECC validation, NCCL bandwidth tests, fabric error sweep and a signed acceptance report you can hand to your own auditors.

Talk to an AI infrastructure architect

Own the GPUs. Let us run them.

Buy your NVIDIA servers from any OEM or distributor you like, ship them to a Velyrix hall, and we handle the rest — deployment, fabric, cooling, monitoring and support. Or rent ours. Either way, you get a plan in one business day.

Frequently asked questions

How is Velyrix GPU Cloud billed?

On-demand instances bill per second with no minimum term and no data egress charges on standard plans. Reserved clusters bill monthly at a fixed contracted rate for terms from one to thirty-six months. Enterprise private clouds are quoted as a fixed monthly platform fee plus capacity.

Do I get bare metal or virtual machines?

Both are available. The default for multi-node training is single-tenant bare metal with direct access to the GPUs, NICs and NVMe. Virtualised instances with GPU passthrough are offered where snapshotting and rapid re-imaging matter more than the last few percent of performance.

What network do multi-node training clusters use?

Rail-optimised non-blocking InfiniBand - NVIDIA Quantum-2 NDR at 400G or Quantum-X800 XDR at 800G per GPU - with SHARP in-network reduction. NVIDIA Spectrum-X Ethernet with RoCEv2 is available where an Ethernet-only operating model is required.

Can I run Slurm and Kubernetes on the same cluster?

Yes. Velyrix commonly partitions a reserved cluster so that a Slurm partition handles long training runs while a Kubernetes partition serves inference, with a shared parallel file system across both. Partition sizes can be adjusted during the term.

Is there a minimum commitment for large clusters?

Clusters above 256 GPUs are normally contracted for a minimum of three months because of fabric build and cabling work. Shorter windows can be accommodated in halls where a matching pod is already built and idle.