Researchers are actively improving memory for LLMs/VLMs. But are we measuring it right?
A system can score 100% by rereading the input for each question. If a human did that, we’d call it terrible memory!
We test beyond accuracy with ECC: Efficiency, Compression, Calibration. 🧵