Long-context LLM inference has a memory problem. As context lengths grow from 128K to over one million tokens, KV cache demands exhaust GPU memory, forcing batch size reductions that severely constrain throughput. Marvell and @SKhynix built a system to address this directly. The CMM-Ax integrates the Marvell Structera A near-memory processing engine with SK hynix DDR5 memory and a co-optimized software stack, enabling KV cache overflow to be offloaded to high-capacity CXL memory and processed where it resides. Testing against standard GPU baselines using the Llama3-8B-1048K model showed up to 5.5 times higher throughput than a single-GPU baseline and 3.6 times higher than a dual-GPU setup. Associate Vice President Khurram Malik and SK hynix Vice President of System Architecture Kangkyu Park detail the architecture and the results: mrvl.co/465xcxv

Last edited Aug 17, 2026 · 5:33 PM UTC

2
24
177
40,787
Sort replies: Relevant Recent Liked
@grok kv cache resides on hbm, correct? What is cvx memory, how does it compare with hbm?
2
2
427