I have been using
@VictorTaelin's solution for memory (OptMem) for 5 weeks now, replacing Claude Code's native memory, and since here at LangWatch we track everything more than anyone, I ran a full analysis on it
I compared OptMem with the 3 months I had before working only with Claude Code Native memory, and luckily, in many of those sessions, Claude randomly didn't activate optmem at all so I kinda had an A/B test done for me
tl;dr:
1. is it better than native? yes, but not by a lot, and in different ways (table below)
2. is memory useful to agents at all? a bit, mostly in long sessions
3. where does it help the most? avoiding rediscoveries and "traps", that is, mistakes that waste time
4. does it save on token cost? meh, a bit
5. are you continuing with optmem? yes
so if you don't know how optmem works, check the gif below and their repo, it has an interesting approach of always keeping the memory a fixed size
this alone is a for me a huge advantage, first and foremost because it gives me such a clear mental model and clear understanding of how my memories are being organized, it's both very structured and simply human readable, not free-for-all random markdown files, nor byte-compressed vector dbs or whatever
second, it limits the damage radius of memory rot. Rot is a fact of life for anything AI agents touches which a human is not supervising. It's like vibe-coding an app through pure bruteforce, you can totally feel the slop compounding and everything rotting. Same happens with your memories, and the wrong ones actually may degrade your harness over time
now not necessarily optmem is a better solution against rot, in fact because of pushing the append-only log it might be worse as agents seem to by default not use the forget functionality a lot (I'm investigating why), or as claude puts it:
"How correction actually happens, per store.
Native rewrites in place. 1,719 of the 2,081 write events are rewrites of an existing file, and MEMORY.md alone was rewritten 601 times. When an agent trips over a wrong file it usually fixes it, as in the thresholds file it rewrote with re-measured numbers. What nobody does is sweep: 39 percent of path tokens went stale when ADR-076 moved the tree, 68 of 260 files fell out of the index, and transcripts show live ls failures on dead index entries. So native rot is silent staleness from the world moving, and it is only corrected when tripped over. Nothing is ever lost or contaminated, and nothing ever stops growing.
OptMem cannot rewrite, so it appends. 281 lines hold 3 explicit correction lines, one of which points at the wrong id, and 5 contradicting pairs with no marker at all. Both halves stay in the log and the tree does not know which one is the correction. That is how the refuted note ended up in wake and the right one did not. So OptMem rot is contradiction plus compression, and nothing corrects it in place by design."
still however, somehow optmem has accumulated less misleading memories than native
there are a lot of low hanging fruits I identified from this which will make optmem better for me, main one being waking it up also on subagents and considering loading it right from a CLAUDE.md import so agents adhere to it more often
I'm also having a claude cleaning up my memory to remove rot, and plan to keep doing this analysis every month or so to keep the house in order
there are a lot of other solutions out there, but memory is hard, and I cant only trust anything if I test it for months like I'm doing here, so that's why I'm favoring simplicity, control, and my own understanding over anything, so I'm sticking to optmem for now