Home/AI Deployment

AI infrastructure deployment & commissioning

From pallets on the dock to a validated cluster

Velyrix deployment engineers build AI factories for operators, enterprises and governments — in our halls or yours. Every engagement ends the same way: a burned-in, benchmarked, documented cluster and a signed acceptance report your auditors can read.

Greenfield build
12 – 20 weeks
Expansion pod
3 – 6 weeks
Cluster scale supported
8 – 4,096 GPUs
Burn-in
72 hours minimum
Handover
Signed acceptance report
Scope

Everything between the purchase order and the first training job

Most GPU projects do not fail on silicon. They fail on coolant loops, on a mis-cabled rail, on firmware drift between eight nodes, or on nobody ever measuring what the fabric actually does.

📐

Design & engineering

Cluster topology, rack elevations, power schedules, coolant loop design, weight and floor loading, fabric cable plan and bill of materials.

🚚

Logistics & receiving

Freight coordination, customs documentation support, secure receiving, asset capture, serial registration and de-trash.

🔨

Rack & stack

Mechanical install of HGX, DGX, MGX and ORv3 systems, rail kits, CDU and manifold plumbing, torque-controlled fastening and ESD-controlled handling.

Power & cooling

Busway and whip termination by licensed electricians, A/B feed balancing, coolant fill and pressure test, leak-detection commissioning.

📡

Fabric build

Rail-optimised InfiniBand or Spectrum-X cabling, labelling to plan, optics seating, link training, error-rate sweep and topology verification.

💾

Firmware baselining

BIOS, BMC, CPLD, NIC, NVSwitch and GPU VBIOS aligned to a single validated baseline across every node, with drift reporting.

💻

Software stack

OS image, NVIDIA driver, CUDA, Fabric Manager, DCGM, container runtime, Slurm or Kubernetes, storage clients and monitoring agents.

🔥

Burn-in & validation

72-hour sustained stress, thermal soak, ECC and XID monitoring, memory row-remap checks, storage and fabric throughput tests.

📋

Acceptance & handover

Documented acceptance test results, as-built drawings, asset register, runbooks, escalation matrix and operator training.

Programme

The Velyrix deployment method

A repeatable nine-stage programme with a named project manager, weekly reporting and a documented gate at the end of every stage.

Discovery & design review

Workload profile, model sizes, token budgets and growth curve translated into a cluster topology, power envelope and storage sizing. Deliverable: design document and BOM.

Site readiness assessment

Power capacity, floor loading, coolant availability, delivery access, fire suppression and fabric pathways surveyed against the design. Deliverable: readiness report with gap list.

Procurement & logistics support

OEM lead-time tracking, staging plan, export and import documentation support, and end-user screening before any controlled hardware ships. Deliverable: delivery schedule.

Mechanical installation

Racks levelled and bonded, systems installed, manifolds and cold plates plumbed, coolant filled and pressure tested, cable management dressed to plan. Deliverable: as-built elevations.

Electrical commissioning

Circuit termination, phase balancing, breaker coordination and metering verification by licensed electricians working to documented method statements. Deliverable: electrical sign-off.

Fabric build & verification

Rail-optimised cabling installed and labelled, every link trained at full rate, bit-error-rate sweep, topology map compared against design. Deliverable: link and topology report.

Software & cluster bring-up

Golden image deployment, firmware baseline enforcement, scheduler configuration, storage mounts, monitoring and alerting integration. Deliverable: configuration baseline.

Burn-in & performance validation

72-hour stress at full power, NCCL all-reduce and all-gather benchmarks, storage throughput, and a reference model training run at target scale. Deliverable: benchmark results.

Acceptance & operational handover

Acceptance test report reviewed and signed, runbooks and escalation paths handed over, operator training delivered, optional transition to Velyrix managed operations. Deliverable: signed ATP.

Acceptance test protocol

What we measure before we call it done

Pass criteria are agreed in writing at design stage and tested at handover. Anything that fails is remediated and re-tested at Velyrix's cost under a fixed-price engagement.

TestMethodTypical pass criterionEvidence
GPU inventory & healthnvidia-smi, DCGM diagnostics level 3100% of GPUs enumerated, zero uncorrectable ECC, zero pending row remapsPer-node DCGM report
Sustained burn-in72 h combined GPU, CPU, memory and NVMe stressZero XID errors, zero thermal throttle events, stable clocksTime-series telemetry export
NVLink / NVSwitchnvbandwidth, p2p bandwidth latency testWithin 5% of platform reference bandwidth on every pairBandwidth matrix
Fabric link qualityLink training, BER sweep, port error countersAll links at full rate, symbol error rate below vendor threshold, zero flaps in 72 hFabric health report
Collective performanceNCCL all-reduce and all-gather, 8 MB – 8 GBBus bandwidth within 10% of topology reference at full scalenccl-tests output
Storage throughputfio and IOR against parallel file systemContracted aggregate read/write and checkpoint time metBenchmark log
Reference training runLlama-class reference model at target node countContracted model FLOPs utilisation and step time achievedTraining log & MFU calculation
Failover & resilienceA/B power pull, leaf switch pull, CDU pump failoverNo job loss on redundant paths, documented recovery timeTest witness sheet
Thermal & powerFull-load soak with inlet, outlet and coolant telemetryAll components within OEM specification at design ambientThermal survey
Security baselineSecure boot, BMC hardening, credential rotation, network segmentationBaseline enforced on 100% of nodes, no default credentialsConfiguration audit
How we work safely

Built to survive an OEM and insurer audit

Deployment work touches other people's warranties, other people's halls and live electrical plant. Velyrix runs the engagement accordingly.

  • Method statements & risk assessments issued and approved before any work on a live floor
  • Licensed electricians for all power termination, working to local code and arc-flash practice
  • ESD-controlled handling aligned to ANSI/ESD S20.20 to protect OEM warranty conditions
  • OEM installation guidance followed so manufacturer warranty and support entitlements remain intact
  • Chain of custody from dock to rack, with serial-level asset capture and photographic records
  • Change control with documented work windows, rollback plans and customer approval gates
  • Insured and screened engineers, background-checked, with per-site access authorisation
  • Export control screening completed before controlled hardware is shipped or installed

Engagement models

Fixed-price build

Agreed scope, agreed acceptance criteria, agreed price. The most common model for greenfield clusters.

Time & materials

Engineer day rates for expansions, remediation, migrations and troubleshooting of an existing estate.

Residency

Velyrix engineers embedded on your site for a defined period to build capability alongside your team.

Build & operate

We build it, then run it under a managed service with an availability and performance SLA.

Typical fixed-price deployment fees run from roughly $2,300 per node for a straightforward air-cooled expansion to $4,600–$10,900 per node for liquid-cooled rack-scale systems including fabric build, validation and documentation. Every engagement is quoted after the design review.

For OEMs, distributors & resellers

We are also the field services bench behind other people's deals

We are set up to deliver deployment work for hardware vendors and their partners — under our brand or, more usefully for you, under theirs.

  • White-label delivery — your brand on the project plan and the acceptance report
  • Validated playbooks for Dell, Supermicro, GIGABYTE, HPE, Lenovo and NVIDIA DGX platforms
  • L1–L3 support with a documented escalation path into the manufacturer
  • RMA coordination and spares depots sized to the deployed fleet
  • Deal registration and a written channel-neutrality policy

Warranty-preserving by construction

Every platform we deploy has a Velyrix playbook derived from the manufacturer's published installation guidance.

  • ESD handling aligned to ANSI/ESD S20.20
  • Torque specifications and rail-kit procedures followed
  • Firmware baselines matched to the vendor's validated stack
  • Serial-level records and photographic evidence retained
  • Failure evidence packaged in the vendor's required format

Result: manufacturer warranty and support entitlements stay intact, and RMA disputes about installation practice do not happen.

Also available

Deployment services beyond the first build

Migration & relocation

Move a live GPU estate between halls or providers with a staged cutover plan, minimal training downtime and full asset reconciliation.

🔍

Cluster health audit

Independent assessment of an existing cluster: fabric errors, firmware drift, thermal headroom, scheduler efficiency and realised MFU, with a prioritised remediation plan.

🔁

Lifecycle & refresh

Capacity planning, staged hardware refresh, decommissioning with NIST SP 800-88 media sanitisation and certified disposal or resale support.

👥

Managed operations

24×7 monitoring, incident response, firmware and driver lifecycle, spares management and capacity reporting against your SLA.

Talk to an AI infrastructure architect

Own the GPUs. Let us run them.

Buy your NVIDIA servers from any OEM or distributor you like, ship them to a Velyrix hall, and we handle the rest — deployment, fabric, cooling, monitoring and support. Or rent ours. Either way, you get a plan in one business day.

Frequently asked questions

What does AI infrastructure commissioning actually include?

Commissioning is the formal process of proving a cluster works to specification before it is handed over: firmware baselining, a 72-hour burn-in, GPU and NVLink health checks, fabric link-quality sweeps, NCCL collective benchmarks, storage throughput tests, a reference training run, failover testing and a thermal survey - all captured in a signed acceptance test report.

Can Velyrix deploy in our own data centre rather than yours?

Yes. A large part of what we do is on customer or third-party sites - enterprise data centres, government facilities and other operators' colocation halls. We work to the host site's access, safety and change-control rules, and we can operate as a subcontractor to your existing integrator if that is simpler.

How long does a greenfield GPU cluster take to deliver?

Once hardware is on site, a typical build runs three to six weeks per pod. End to end - design, site works, hardware lead time, build and commissioning - a greenfield AI factory is normally 12 to 20 weeks, with OEM lead time the largest variable.

Will third-party deployment work void our OEM hardware warranty?

No, when the work follows the manufacturer's published installation guidance. Velyrix engineers work to OEM procedures with ESD-controlled handling, torque specifications and documented serial-level records, so manufacturer warranty and support entitlements remain intact.

What happens if the cluster fails an acceptance test?

Under a fixed-price engagement, Velyrix remediates and re-tests at our cost until the agreed criteria are met. Failures traced to defective hardware are raised with the OEM under warranty, and we manage that process on your behalf.

Do you support liquid-cooled and rack-scale systems such as GB300 NVL72?

Yes. Liquid-cooled deployments include manifold and cold-plate installation, coolant fill and pressure testing, CDU commissioning, leak-detection validation and pump failover testing, in addition to the standard compute and fabric acceptance protocol.