It’ll be interesting to see how open and closed-weight models diverge on memory.
With open weights, you can try creative approaches like Doc-to-LoRA or Code-to-LoRA where instead of stuffing the same documents into context, we compile some of that knowledge into the model itself.
Closed labs can do this privately too, of course. But as a user, all you'll typically see from them is bigger context and more compute.
The funny outcome may be that the less elegant approach still wins.
My hot take is that markdown files are a horrible way to structure memory.
And that's why agents make so many stupid mistakes.
They perform the same way a human amnesia and having to read a bunch of docs before doing a basic task would: poorly, cutting corners.
Intelligence is not more data.
Indexing docs in a graph (presumptuously called knolwdge graph) is just better search. It doesn't solve the fundamental problem.