Two DGX Station GB300 towers (an ASUS ExpertCenter Pro ET900N G3 and an HP ZGX Fury AI Station) were linked with two 400G direct-attach copper cables between their ConnectX-8 QSFP112 ports and configured with NVIDIA’s playbook into two point-to-point RoCEv2 rails. The ConnectX-8 also serves as a secondary I/O hub for Grace and PCIe Gen6 storage. Low-level tests hit near line rate: NCCL collective broadcasts, reduces, and send/receive reached about 98 GB/s across the pair; GPUDirect RDMA hit 392 Gb/s per rail (98% of 400G); host RDMA from Grace reached ~228 Gb/s; and 8-byte RDMA writes averaged 1.46 µs. The guide setup was straightforward and validated both rails, jumbo MTU, and NICs for NCCL.
Benchmarking focused on inference with tensor parallelism and prefill/decode disaggregation. For models that overflow a single station’s HBM, two Stations delivered massive gains: GLM-5.2 (433 GB) rose from 139 to 1,525 output tokens/sec at 32 streams (≈11×) and to 3,141 at 256; GLM-5.3 jumped from 188 to 5,018 tps (≈27×) and sustained ~600 tps on 8,192-token prompts (≈5×). Minimax-M3 long prompts improved 7.1× at eight streams and scaled to 1,801 tps at 256. For models that fit on a single GPU, running a second Station as a replica and splitting prefill/decode tripled or doubled throughput in many cases (GLM-5.3 Flash and DeepSeek V4.1 Flash reached multi-thousand token/s rates), showing that adding a second GB300 can either expand model capacity or substantially increase throughput depending on workload.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.