Skip to content
GPU server maintenance for dense AI racks.
GPU Maintenance

GPU server maintenance for dense AI racks.

NVIDIA DGX and HGX, Dell PowerEdge XE, HPE Apollo, Supermicro, and Lenovo GPU nodes — break/fix, parts, and rack deployments. Same contract as the rest of the floor.

What’s included

The GPU, the node, and the fabric next to it.

GPU racks run hotter, pull more power, and sit under jobs that cannot wait on a vendor spare. NETRAID covers the hardware — current Hopper and Blackwell, and the A100 rows that are already off warranty.

01

GPU node break/fix

SXM and PCIe GPUs, NVSwitch, NVLink bridges, baseboards, PSUs, and fans. The part that failed — not a wholesale swap of the rack.

02

Dense rack deployments

GPU-dense rows take more power, more copper, and a tighter install window. We rack, cable, and hand off the nodes so training jobs can start.

03

InfiniBand and Spectrum fabric

The cluster is only as good as the switch it hangs off. QM and Spectrum hardware sits on the same contract as the DGX and XE nodes.

04

Air and liquid-cooled systems

Direct-to-chip and rear-door loops are in scope for the server hardware. We replace the node, the GPU, and the cold plate — not the building CDU.

05

Same contract as the rest of the rack

GPU nodes next to PowerEdge, ProLiant, and the SAN. One SLA mix, one invoice, one Client Portal.

06

Parts that match the density

GPU trays, NVSwitch boards, and high-watt PSUs staged near the site. 24×7×2 on the training cluster, next-business-day on the lab.

How it works

How GPU coverage starts.

  1. 01

    The cluster, not a serial

    How many nodes, which GPU, air or liquid, and whether InfiniBand sits next to them. A short inventory is enough to start.

  2. 02

    Coverage that matches the job

    Training clusters can be 24×7×2. Inference and lab stay on a slower window. Mix them on one contract.

  3. 03

    Parts already nearby

    GPU, NVSwitch, and PSU staged before the first ticket. When a node drops, the part is not a week out.

FAQ

Questions About GPU Coverage.

Do you replace the GPU itself, or only the server around it?

Both. SXM and PCIe GPUs, NVSwitch, NVLink bridges, and the host — fans, PSUs, baseboards, and cables.

Can GPU nodes sit on the same contract as the rest of the rack?

Yes. Dell, HPE, Supermicro, Lenovo, NVIDIA DGX, and the InfiniBand fabric can share one SLA mix and one invoice.

Do you cover liquid-cooled GPU servers?

Yes — the node, GPU, and cold plate in air or liquid-cooled racks. Building CDUs and facility water are out of scope unless we quote them separately.

Will you help stand up a new GPU row?

Yes. Rack-and-stack, cabling, and the fabric next to the nodes. Tell us the floor, the power, and the SKU. We return a scoped quote.

What about DGX and HGX after the vendor warranty?

That is the usual ask. Current Blackwell and Hopper, and older A100 DGX that still trains. If the serial is not listed, send it.

Same-day quote

Request GPU coverage.

Node count, GPU type, air or liquid, and whether the fabric is in scope. We return a scoped quote — break/fix, parts, and deployments.

Hardware Maintenance Quote

Select every type this quote should cover.

Get a GPU maintenance quote.

Tell us the nodes and the GPU. We return a scoped quote.