This really matters for Local AI! We keep shrinking the model weights. Now researchers are going after the model's working memory. A new technique called MILO cuts KV-cache memory by up to ๐Ÿ‘‰ 50% ๐Ÿ‘ˆ and reports up to 1.8ร— higher throughput on Qwen2.5 experiments. ๐ŸŽฏ This is really signficant because KV cache grows with context. So once the quantized model fits into VRAM/RAM, context is the next memory hog. As an example ... ๐Ÿง  8GB KV cache @ 128K context โ†“ 50% compression ๐Ÿง  ~4GB for the same cached context So that saved memory could mean ... ๐Ÿ“š ~2ร— the context for the same KV-memory budget ๐Ÿค– more simultaneous agents/sessions ๐Ÿ’พ less KV spilling into slower RAM ๐Ÿง  more room for model weights MILO breaks the KV cache into blocks, finds redundancy using low-rank compression, and gives information-dense blocks more precision while compressing redundant blocks harder. โš ๏ธ MILO currently targets many-shot long-context inference and was tested on Qwen2.5 3B/7B with an Nvidia L4. Obviously this is research and not a llama.cpp feature you can turn on tonight. ๐Ÿ”— Paper in ALT

Sep 26, 2026 ยท 4:00 AM UTC

12
20
87
3,143
Sort replies: Relevant Recent Liked
Replying to @TeksEdge
Nice! these papers are things that need to be fed to frontier models, not stardew valley or grand theft auto clones.
56
Replying to @TeksEdge
whatever this turns into in production, the constraint to design around is locality: many-shot is max prefix-reuse territory, and reuse only keeps paying if a shared prefix encodes identically run to run. block-local compression keeps both the cache and the memory win; context-aware trades cache hits for bytes.
56