01
GPU node break/fix
SXM and PCIe GPUs, NVSwitch, NVLink bridges, baseboards, PSUs, and fans. The part that failed — not a wholesale swap of the rack.

NVIDIA DGX and HGX, Dell PowerEdge XE, HPE Apollo, Supermicro, and Lenovo GPU nodes — break/fix, parts, and rack deployments. Same contract as the rest of the floor.
What’s included
GPU racks run hotter, pull more power, and sit under jobs that cannot wait on a vendor spare. NETRAID covers the hardware — current Hopper and Blackwell, and the A100 rows that are already off warranty.
01
SXM and PCIe GPUs, NVSwitch, NVLink bridges, baseboards, PSUs, and fans. The part that failed — not a wholesale swap of the rack.
02
GPU-dense rows take more power, more copper, and a tighter install window. We rack, cable, and hand off the nodes so training jobs can start.
03
The cluster is only as good as the switch it hangs off. QM and Spectrum hardware sits on the same contract as the DGX and XE nodes.
04
Direct-to-chip and rear-door loops are in scope for the server hardware. We replace the node, the GPU, and the cold plate — not the building CDU.
05
GPU nodes next to PowerEdge, ProLiant, and the SAN. One SLA mix, one invoice, one Client Portal.
06
GPU trays, NVSwitch boards, and high-watt PSUs staged near the site. 24×7×2 on the training cluster, next-business-day on the lab.
Platforms
If the badge is not listed, send the serial. We still quote it.
How it works
01
How many nodes, which GPU, air or liquid, and whether InfiniBand sits next to them. A short inventory is enough to start.
02
Training clusters can be 24×7×2. Inference and lab stay on a slower window. Mix them on one contract.
03
GPU, NVSwitch, and PSU staged before the first ticket. When a node drops, the part is not a week out.
FAQ
Both. SXM and PCIe GPUs, NVSwitch, NVLink bridges, and the host — fans, PSUs, baseboards, and cables.
Yes. Dell, HPE, Supermicro, Lenovo, NVIDIA DGX, and the InfiniBand fabric can share one SLA mix and one invoice.
Yes — the node, GPU, and cold plate in air or liquid-cooled racks. Building CDUs and facility water are out of scope unless we quote them separately.
Yes. Rack-and-stack, cabling, and the fabric next to the nodes. Tell us the floor, the power, and the SKU. We return a scoped quote.
That is the usual ask. Current Blackwell and Hopper, and older A100 DGX that still trains. If the serial is not listed, send it.
Same-day quote
Node count, GPU type, air or liquid, and whether the fabric is in scope. We return a scoped quote — break/fix, parts, and deployments.

Tell us the nodes and the GPU. We return a scoped quote.