Bottlenecks never disappear; they just shift to another layer of the stack.
Before I came to the AI Infra Summit, I was biased into thinking that shifting to the optical domain to speed up scale-up and memory BW will completely solve the problem of low compute utilization by moving data faster.
However, after sitting through keynotes and expert panels in the data movement track, I realize that I failed to appreciate the system-level complexity that goes into moving and orchestrating massive amounts of data from concurrent users.
I realized how "improvements" tend to add additional complexity burden and move bottlenecks to other parts of the system. More scale-up and memory BW helps to push performance, but isn’t the only means to improve overall cluster data throughput.
Next week, I'll be releasing my 3-part post that explores how a token goes through a network and the challenges with scaling performance and increasing utilization.
It synthesizes from a wealth of knowledge from domain-specific experts from major players such as Credo, Astera, AMD, Microsoft, Broadcom, Marvell, as well as emerging companies like UpscaleAI, iPronics, and Salience Labs.
Make sure you follow and subscribe to my Substack, "Silicon-Co-design" so you'll be the first to know about it.