Kylie Zhang retweeted
Researchers are actively improving memory for LLMs/VLMs. But are we measuring it right? A system can score 100% by rereading the input for each question. If a human did that, we’d call it terrible memory! We test beyond accuracy with ECC: Efficiency, Compression, Calibration. 🧵
4
7
22
960
Kylie Zhang retweeted
Love this. Makes you take a step back and appreciate how far we've come—but also how much more can be achieved. ordinaryabundance.com/
1
13
1,022