Agent Memory Leaderboard Official. First open championship for agent memory systems. Neutral · Reproducible · Fair. mailto: contactus@agentmemoryleaderboard.ai

Based in United States
Agent Memory Challenge 2026 Cycle 2 opens September 20. A shared evaluation for long-term Agent Memory across Textual, Coding, and Multimodal tracks—testing not just what agents store, but whether they retrieve current, useful evidence when it matters.
Article

From “Storing” to “Staying Current”: Why Agent Memory Needs a Shared Evaluation

Agent Memory Challenge 2026 Cycle 2 opens September 20 across Textual, Coding, and Multimodal Memory. Long-term Agent Memory is not about keeping more history. It is about retrieving the right

6
1
9
694
Can past code, bugs, and development trails help an Agent solve the next task? AML’s first Coding Memory results compare industry and open-source approaches—and show why memory is moving from storing history to reducing repeated trial and error.
Article

From “Remembering Code” to “Solving Tasks”: How Coding Memory Helps Agents Reuse Engineering Experie

From “Remembering Code” to “Solving Tasks”: How Coding Memory Helps Agents Reuse Engineering Experience Compared with textual memory, Coding Memory asks a more direct question: Can experience

5
2
12
866
How does a Markdown decision log become Agent Memory? FlowGrid AML Retriever ranked #8 on the first AML Open Leaderboard (43.98). This deep dive explores its evidence-first retrieval, deterministic hybrid search, and path toward governed memory.
Article

#From “Deciding” to “Retrieving”: How FlowGrid Turns Project History into Agent Memory Evidence

FlowGrid AML Retriever ranked #8 on the first AML Open Leaderboard with an Overall Score of 43.98. But what makes its approach interesting? Many Agent Memory systems begin with a familiar question:

9
1
11
1,014
AML Technical Deep Dive #3: ChronoHybridMem ranked #5 with 44.33. Its core idea is simple: Agent Memory should return evidence, not just similar text. Raw messages, source-linked facts, dual-path retrieval, and constrained reranking make memory verifiable.
Article

From “Retrieving” to “Verifying”: How ChronoHybridMem Turns Agent Memory into Evidence

ChronoHybridMem ranked #5 on the first AML Open Leaderboard with an Overall Score of 44.33. But what makes its approach interesting? Long-term memory retrieval for an AI agent is often framed as a

3
1
8
645
What if a memory system didn’t try to manage memory at all? ActiveMemoryIndex ranked #3 on the first AML Open Leaderboard by preserving raw evidence, minimizing processing, and letting the model decide. Here’s how it works ↓
Article

From “Managing” to “Preserving”: How ActiveMemoryIndex Keeps Memory Simple

How ActiveMemoryIndex Ranks #3 Without Memory Governance AI memory systems often try to make memory smarter. They summarize conversations, extract facts, resolve conflicts, update old values, build

2
6
350
What if AI agents didn’t need to remember everything? ReFind ranked #2 on the first AML Open Leaderboard (44.97). Our latest deep dive explores how it lets agents actively search raw chat history — and where query-based retrieval reaches its limits.
Article

From “Remembering” to “Searching”: How ReFind Lets Agents Explore Raw Chat History

ReFind ranked #2 on the first AML Open Leaderboard with a score of 44.97. But what makes its approach interesting? Retrieving memory for an AI agent sounds simple: Given the current query, find the

2
8
539
Similarity isn’t enough for agent memory. InvMem ranked #1 on the first AML Open Leaderboard (45.06). From Dense + BM25 to Weighted RRF and context expansion, here’s how it turns relevant memories into complete evidence. 👇
Article

AML Technical Deep Dive #1: From “Similarity” to “Completeness” — How InvMem Retrieves Useful Long-T

InvMem ranked #1 on the first AML Open Leaderboard with a score of 45.06. But what makes its approach interesting? Retrieving memory for an AI agent sounds simple: Given the current query, find the

4
7
208
The first season of Agent Memory Leaderboard (AML) is complete. Explore how 100+ teams, 10+ benchmarks, and different memory approaches were evaluated under a unified framework — and what this reveals about the future of AI agent memory.
Article

Agent Memory Leaderboard First Season Recap: Toward a Standardized Evaluation Framework for AI Memor

Building a common language for AI memory evaluation As AI agents move beyond short conversations into long-horizon tasks such as coding, research, and personalized assistance, memory has become one of

4
12
5,405