Is it just me or are existing agent memory datasets underwhelming?
They to some extend assume retrieval will just work, they tell you roughly where the right answer is, and then they sit back and see if you can piece together a response.
This doesn't resemble what the work feels like in my experience...