This really matters for Local AI! We keep shrinking the model weights.
Now researchers are going after the model's working memory.
A new technique called MILO cuts KV-cache memory by up to ๐ 50% ๐ and reports up to 1.8ร higher throughput on Qwen2.5 experiments.
๐ฏ This is really signficant because KV cache grows with context.
So once the quantized model fits into VRAM/RAM, context is the next memory hog.
As an example ...
๐ง 8GB KV cache @ 128K context
โ 50% compression
๐ง ~4GB for the same cached context
So that saved memory could mean ...
๐ ~2ร the context for the same KV-memory budget
๐ค more simultaneous agents/sessions
๐พ less KV spilling into slower RAM
๐ง more room for model weights
MILO breaks the KV cache into blocks, finds redundancy using low-rank compression, and gives information-dense blocks more precision while compressing redundant blocks harder.
โ ๏ธ MILO currently targets many-shot long-context inference and was tested on Qwen2.5 3B/7B with an Nvidia L4.
Obviously this is research and not a llama.cpp feature you can turn on tonight.
๐ Paper in ALT
Sep 26, 2026 ยท 4:00 AM UTC
12
20
87
3,143



