We were the first cloud provider to bring up and validate
@NVIDIA Vera Rubin NVL72. The underlying challenge is making several racks behave as one cluster.
That means automated rack lifecycle (Racky, Valvey, Rack LifeCycle Controller), straggler detection down to a single slow GPU, and a two-tier, non-blocking RoCE fabric at 1.6 Tb/s per GPU, designed to reach ~128,000 GPUs.
How we did it:
utm.io/usUQn