Reduced GPU capacity in US-CENTRAL-08A (DH204)
- update
August 20, 2026 5:12PM UTC Investigating - We are investigating a loss of production GPU capacity affecting a subset (19) of GB300 racks in DH204. Affected nodes have intermittently gone unhealthy and been evicted; a number have been taken out of production for recovery. Affected: GB300 compute Current impact: Several racks are below the 16-node threshold for a full NVL64 domain, so multi-node jobs targeting those racks may fail to place or may be interrupted. Capacity in other datahalls and regions is unaffected. The first node losses were observed at approximately 2026-08-19 19:00 UTC. Fleet Operations has been recovering nodes by reboot since then, and additional engineering teams were engaged at 16:30 UTC today. We will keep you posted on further details as we progress. Thank you for your patience and cooperation while we are investigating.