Home/AI Deployment
AI infrastructure deployment & commissioningFrom pallets on the dock to a validated cluster
Velyrix deployment engineers build AI factories for operators, enterprises and governments — in our halls or yours. Every engagement ends the same way: a burned-in, benchmarked, documented cluster and a signed acceptance report your auditors can read.
Everything between the purchase order and the first training job
Most GPU projects do not fail on silicon. They fail on coolant loops, on a mis-cabled rail, on firmware drift between eight nodes, or on nobody ever measuring what the fabric actually does.
Design & engineering
Cluster topology, rack elevations, power schedules, coolant loop design, weight and floor loading, fabric cable plan and bill of materials.
Logistics & receiving
Freight coordination, customs documentation support, secure receiving, asset capture, serial registration and de-trash.
Rack & stack
Mechanical install of HGX, DGX, MGX and ORv3 systems, rail kits, CDU and manifold plumbing, torque-controlled fastening and ESD-controlled handling.
Power & cooling
Busway and whip termination by licensed electricians, A/B feed balancing, coolant fill and pressure test, leak-detection commissioning.
Fabric build
Rail-optimised InfiniBand or Spectrum-X cabling, labelling to plan, optics seating, link training, error-rate sweep and topology verification.
Firmware baselining
BIOS, BMC, CPLD, NIC, NVSwitch and GPU VBIOS aligned to a single validated baseline across every node, with drift reporting.
Software stack
OS image, NVIDIA driver, CUDA, Fabric Manager, DCGM, container runtime, Slurm or Kubernetes, storage clients and monitoring agents.
Burn-in & validation
72-hour sustained stress, thermal soak, ECC and XID monitoring, memory row-remap checks, storage and fabric throughput tests.
Acceptance & handover
Documented acceptance test results, as-built drawings, asset register, runbooks, escalation matrix and operator training.
The Velyrix deployment method
A repeatable nine-stage programme with a named project manager, weekly reporting and a documented gate at the end of every stage.
Discovery & design review
Workload profile, model sizes, token budgets and growth curve translated into a cluster topology, power envelope and storage sizing. Deliverable: design document and BOM.
Site readiness assessment
Power capacity, floor loading, coolant availability, delivery access, fire suppression and fabric pathways surveyed against the design. Deliverable: readiness report with gap list.
Procurement & logistics support
OEM lead-time tracking, staging plan, export and import documentation support, and end-user screening before any controlled hardware ships. Deliverable: delivery schedule.
Mechanical installation
Racks levelled and bonded, systems installed, manifolds and cold plates plumbed, coolant filled and pressure tested, cable management dressed to plan. Deliverable: as-built elevations.
Electrical commissioning
Circuit termination, phase balancing, breaker coordination and metering verification by licensed electricians working to documented method statements. Deliverable: electrical sign-off.
Fabric build & verification
Rail-optimised cabling installed and labelled, every link trained at full rate, bit-error-rate sweep, topology map compared against design. Deliverable: link and topology report.
Software & cluster bring-up
Golden image deployment, firmware baseline enforcement, scheduler configuration, storage mounts, monitoring and alerting integration. Deliverable: configuration baseline.
Burn-in & performance validation
72-hour stress at full power, NCCL all-reduce and all-gather benchmarks, storage throughput, and a reference model training run at target scale. Deliverable: benchmark results.
Acceptance & operational handover
Acceptance test report reviewed and signed, runbooks and escalation paths handed over, operator training delivered, optional transition to Velyrix managed operations. Deliverable: signed ATP.
What we measure before we call it done
Pass criteria are agreed in writing at design stage and tested at handover. Anything that fails is remediated and re-tested at Velyrix's cost under a fixed-price engagement.
| Test | Method | Typical pass criterion | Evidence |
|---|---|---|---|
| GPU inventory & health | nvidia-smi, DCGM diagnostics level 3 | 100% of GPUs enumerated, zero uncorrectable ECC, zero pending row remaps | Per-node DCGM report |
| Sustained burn-in | 72 h combined GPU, CPU, memory and NVMe stress | Zero XID errors, zero thermal throttle events, stable clocks | Time-series telemetry export |
| NVLink / NVSwitch | nvbandwidth, p2p bandwidth latency test | Within 5% of platform reference bandwidth on every pair | Bandwidth matrix |
| Fabric link quality | Link training, BER sweep, port error counters | All links at full rate, symbol error rate below vendor threshold, zero flaps in 72 h | Fabric health report |
| Collective performance | NCCL all-reduce and all-gather, 8 MB – 8 GB | Bus bandwidth within 10% of topology reference at full scale | nccl-tests output |
| Storage throughput | fio and IOR against parallel file system | Contracted aggregate read/write and checkpoint time met | Benchmark log |
| Reference training run | Llama-class reference model at target node count | Contracted model FLOPs utilisation and step time achieved | Training log & MFU calculation |
| Failover & resilience | A/B power pull, leaf switch pull, CDU pump failover | No job loss on redundant paths, documented recovery time | Test witness sheet |
| Thermal & power | Full-load soak with inlet, outlet and coolant telemetry | All components within OEM specification at design ambient | Thermal survey |
| Security baseline | Secure boot, BMC hardening, credential rotation, network segmentation | Baseline enforced on 100% of nodes, no default credentials | Configuration audit |
Built to survive an OEM and insurer audit
Deployment work touches other people's warranties, other people's halls and live electrical plant. Velyrix runs the engagement accordingly.
- Method statements & risk assessments issued and approved before any work on a live floor
- Licensed electricians for all power termination, working to local code and arc-flash practice
- ESD-controlled handling aligned to ANSI/ESD S20.20 to protect OEM warranty conditions
- OEM installation guidance followed so manufacturer warranty and support entitlements remain intact
- Chain of custody from dock to rack, with serial-level asset capture and photographic records
- Change control with documented work windows, rollback plans and customer approval gates
- Insured and screened engineers, background-checked, with per-site access authorisation
- Export control screening completed before controlled hardware is shipped or installed
Engagement models
Agreed scope, agreed acceptance criteria, agreed price. The most common model for greenfield clusters.
Engineer day rates for expansions, remediation, migrations and troubleshooting of an existing estate.
Velyrix engineers embedded on your site for a defined period to build capability alongside your team.
We build it, then run it under a managed service with an availability and performance SLA.
Typical fixed-price deployment fees run from roughly $2,300 per node for a straightforward air-cooled expansion to $4,600–$10,900 per node for liquid-cooled rack-scale systems including fabric build, validation and documentation. Every engagement is quoted after the design review.
We are also the field services bench behind other people's deals
We are set up to deliver deployment work for hardware vendors and their partners — under our brand or, more usefully for you, under theirs.
- White-label delivery — your brand on the project plan and the acceptance report
- Validated playbooks for Dell, Supermicro, GIGABYTE, HPE, Lenovo and NVIDIA DGX platforms
- L1–L3 support with a documented escalation path into the manufacturer
- RMA coordination and spares depots sized to the deployed fleet
- Deal registration and a written channel-neutrality policy
Warranty-preserving by construction
Every platform we deploy has a Velyrix playbook derived from the manufacturer's published installation guidance.
- ESD handling aligned to ANSI/ESD S20.20
- Torque specifications and rail-kit procedures followed
- Firmware baselines matched to the vendor's validated stack
- Serial-level records and photographic evidence retained
- Failure evidence packaged in the vendor's required format
Result: manufacturer warranty and support entitlements stay intact, and RMA disputes about installation practice do not happen.
Deployment services beyond the first build
Migration & relocation
Move a live GPU estate between halls or providers with a staged cutover plan, minimal training downtime and full asset reconciliation.
Cluster health audit
Independent assessment of an existing cluster: fabric errors, firmware drift, thermal headroom, scheduler efficiency and realised MFU, with a prioritised remediation plan.
Lifecycle & refresh
Capacity planning, staged hardware refresh, decommissioning with NIST SP 800-88 media sanitisation and certified disposal or resale support.
Managed operations
24×7 monitoring, incident response, firmware and driver lifecycle, spares management and capacity reporting against your SLA.
Own the GPUs. Let us run them.
Buy your NVIDIA servers from any OEM or distributor you like, ship them to a Velyrix hall, and we handle the rest — deployment, fabric, cooling, monitoring and support. Or rent ours. Either way, you get a plan in one business day.
Frequently asked questions
What does AI infrastructure commissioning actually include?
Commissioning is the formal process of proving a cluster works to specification before it is handed over: firmware baselining, a 72-hour burn-in, GPU and NVLink health checks, fabric link-quality sweeps, NCCL collective benchmarks, storage throughput tests, a reference training run, failover testing and a thermal survey - all captured in a signed acceptance test report.
Can Velyrix deploy in our own data centre rather than yours?
Yes. A large part of what we do is on customer or third-party sites - enterprise data centres, government facilities and other operators' colocation halls. We work to the host site's access, safety and change-control rules, and we can operate as a subcontractor to your existing integrator if that is simpler.
How long does a greenfield GPU cluster take to deliver?
Once hardware is on site, a typical build runs three to six weeks per pod. End to end - design, site works, hardware lead time, build and commissioning - a greenfield AI factory is normally 12 to 20 weeks, with OEM lead time the largest variable.
Will third-party deployment work void our OEM hardware warranty?
No, when the work follows the manufacturer's published installation guidance. Velyrix engineers work to OEM procedures with ESD-controlled handling, torque specifications and documented serial-level records, so manufacturer warranty and support entitlements remain intact.
What happens if the cluster fails an acceptance test?
Under a fixed-price engagement, Velyrix remediates and re-tests at our cost until the agreed criteria are met. Failures traced to defective hardware are raised with the OEM under warranty, and we manage that process on your behalf.
Do you support liquid-cooled and rack-scale systems such as GB300 NVL72?
Yes. Liquid-cooled deployments include manifold and cold-plate installation, coolant fill and pressure testing, CDU commissioning, leak-detection validation and pump failover testing, in addition to the standard compute and fabric acceptance protocol.