Two
$MRVL EVPs were interviewed to talk about the shift from copper to optics, the memory wall, and their custom silicon strategy. Notes below.
Clip is of Dave Lazovsky (EVP, GM of Data Center Networking) saying scale-up optics on an analog SerDes can cut total data-center power more than 25%.
On with Dave is Will Chu, EVP and GM of the Custom Cloud Solutions Business.
1. The pitch, Marvell is not just an XPU company
The market honed in on the XPU as the one custom element, but Marvell sells "XPU attach," everything around it across networking, memory, storage and security.
For every XPU there are three, four or more attach opportunities in the next-generation design.
2. A concentrated market and a custom business model
This is "the largest infrastructure investment in the history of humanity," and it is very concentrated.
Four companies represent north of 75% of the total addressable market in data-center infrastructure.
3. The forcing function, copper to optics and the memory wall
"AI infrastructure is the forcing function," and the shift from copper to optical scale-up networks "comes down to the models."
Foundation models first drove raw memory capacity, then reasoning models and inference-time compute piled on KV cache.
KV caches grew about 10X in nine months, part of why Micron and Hynix now carry more than $1 trillion in market value.
Housing that HBM capacity forces scaling beyond the rack, past 144 interconnected XPUs toward 576 in a pod.
4. The interconnect portfolio, SerDes plus CPO and NPO
Marvell claims "the world's strongest team" in optical interconnect, backed by 224 gig SerDes moving to 448 gig, integrated onto the XPU.
It deploys CPO directly on its own switches for both Ethernet and UAL scale-up, and offers XPU-side links with both CPO and multiple NPO engagements.
NPO, or onboard optics, eases customers into the first wave of optics before committing to CPO.
Will's addition is that the portfolio walks a customer from copper to NPO to CPO end to end, and putting the same IO on both sides of the scale-up link avoids "link flops."
The optics come as "fast pipe" at 200 gig going to 400 gig and "flat pipe" at 56 gig going to about 112 gig, with the analog SerDes far more energy efficient.
5. Getting over the memory wall, the in-box techniques
Dense SRAM, specialized in-chip IP tightly coupled to compute, since off-the-shelf memory compilers no longer meet customer needs.
Custom HBM for more bandwidth and capacity, with 3D stacking of memory on the logic die coming next.
Memory expansion outside the chip through CXL or proprietary methods to add raw capacity.
Disaggregated memory like the Photonic Fabric Memory Appliance, which puts 32 terabytes in a separate rack.
Near-memory compute, placing CPUs right next to memory to offload the CPU or GPU, since every technique still drives more IO and connectivity.
6. Memory's flip to best-in-industry margins and the CXL moment
Memory went from commoditized, at "multi-digit negative gross profit margins," to "the highest margins in the industry," led by HBM.
High memory prices are pushing hyperscale architects to rethink whole infrastructures.
CXL demand is rising to recycle old memory onto new servers, with compression giving about 2X the memory at 1X the chips, and pooling still to come.
The next bleed-over is into flash, both HBF and traditional flash, to tier memory below DRAM cost.
7. The Photonic Fabric Memory Appliance
The appliance holds 32 terabytes and is built to remove the trade-off between DDR capacity and cost and HBM bandwidth.
Multiple pseudo-channels per HBM stack plus parallel memory transactions hide latency.
For the first time it lets customers scale memory capacity and bandwidth independent of compute.
8. Customization in scale-up networking
No two hyperscalers have standardized, so each builds its own customized scale-up network at both the physical and logical layers.
Some lean toward Ethernet scale-up, some toward UAL (memory-semantic load-store, more like NVLink and more efficient than moving Ethernet packets), and one runs a proprietary topology.
Marvell customizes down to the protocol layer and tunes the FEC per customer, since "this is not a merchant silicon world any longer" and it co-architects switches years ahead.
Will frames the multi-trillion-dollar build as the economic rationale, with the XPU and its scale-up network "very, very tightly coupled."
9. Scale-across, coherent optics, and the power problem
Will's three elements are connectivity, memory and compute, with memory the key to the memory wall and compute the better-understood piece.
Data-center campuses are exploding, and Zuckerberg said Meta's newest is about the size of Manhattan.
A new class of optics for sub-10 kilometer reach, coherent light, has taken off, and Marvell claims it is first to market with 1.6T optical.
Scale-up is roughly 85% of data-center traffic, processor to processor, versus 15% scale-out, so efficiency there drives total power.
Power turns into a burning problem over the next three to five years, with the US "nowhere near enough" to build out, so Marvell's optical scale-up on analog SerDes cuts link power about 4X versus conventional 224 gig optics.
Fully deployed, Marvell believes that lands a 25% plus cut in total data-center power.
Dylan Patel of SemiAnalysis says new PCIe switches ($MRVL
$ALAB) are turning the whole rack into the unit of AI system design.
"Historically, Microchip and Broadcom had PCI switches, but now there's new PCI switches, whether it be Marvell or Astera Labs, 256 lanes, 320 lanes. This enables people to make these really interesting topologies."
"I've seen some Supermicro servers based on these newer switches that enable much larger domains of, whether it be storage, accelerators, NICs, a lot more customization that's available."
"It's not just about one server now, it's about a whole rack."
____
Follow
@firesidealpha for more daily highlights of key business and technology conversations.