Long-context LLM inference has a memory problem. As context lengths grow from 128K to over one million tokens, KV cache demands exhaust GPU memory, forcing batch size reductions that severely constrain throughput.
Marvell and @SKhynix built a system to address this directly. The CMM-Ax integrates the Marvell Structera A near-memory processing engine with SK hynix DDR5 memory and a co-optimized software stack, enabling KV cache overflow to be offloaded to high-capacity CXL memory and processed where it resides. Testing against standard GPU baselines using the Llama3-8B-1048K model showed up to 5.5 times higher throughput than a single-GPU baseline and 3.6 times higher than a dual-GPU setup.
Associate Vice President Khurram Malik and SK hynix Vice President of System Architecture Kangkyu Park detail the architecture and the results: mrvl.co/465xcxv
Last edited Aug 17, 2026 · 5:33 PM UTC
2
24
177
40,787


