Agent Memory Leaderboard Official. First open championship for agent memory systems. Neutral · Reproducible · Fair. mailto: contactus@agentmemoryleaderboard.ai

Agent Memory Challenge 2026 Cycle 2 is now open. Long-term memory is not just about storing more history. It is about retrieving the right evidence, recognizing what has changed, and avoiding stale context when an agent needs to act. Three tracks: Textual · Coding · Multimodal Open-source Methods · Commercial Products Over USD 22,000 prize pool for eligible open-source teams. A shared Add/Search interface. Standardized Answer/Eval. Public, comparable results. Join: agentmemoryleaderboard.ai/ev…
12
7
31
27,029
Shoutout to @omarsar0 for sharing Agent Memory Leaderboard Cycle 2! 🙌 We now have 50+ teams signed up — the event is picking up steam fast. Don’t miss your chance to join the challenge. agentmemories.ai/competition…
Recommended if you work on agent memory. Retrieval is the hardest memory problem in most harnesses because agents keep pulling stale context. AML tests this with a coding track of 150 software tasks, each run with relevant history and again with noisy history.
6
11
448
Agent Memory Challenge 2026 Cycle 2 is now open. Long-term memory is not just about storing more history. It is about retrieving the right evidence, recognizing what has changed, and avoiding stale context when an agent needs to act. Three tracks: Textual · Coding · Multimodal Open-source Methods · Commercial Products Over USD 22,000 prize pool for eligible open-source teams. A shared Add/Search interface. Standardized Answer/Eval. Public, comparable results. Join: agentmemoryleaderboard.ai/ev…
12
7
31
27,029
For the full technical overview of Cycle 2—including the shared Add/Search evaluation boundary, Textual, Coding, and Multimodal tracks, and the principles behind reproducible Agent Memory evaluation:
Agent Memory Challenge 2026 Cycle 2 opens September 20. A shared evaluation for long-term Agent Memory across Textual, Coding, and Multimodal tracks—testing not just what agents store, but whether they retrieve current, useful evidence when it matters.
Article

From “Storing” to “Staying Current”: Why Agent Memory Needs a Shared Evaluation

Agent Memory Challenge 2026 Cycle 2 opens September 20 across Textual, Coding, and Multimodal Memory. Long-term Agent Memory is not about keeping more history. It is about retrieving the right

3
264
Read the participation guide: agentmemoryleaderboard.ai/ru… API integration guide: agentmemoryleaderboard.ai/ap… Open-source evaluation framework and updates: github.com/AML-memory/agent-… Cycle 2 runs through October 31. Official results are planned for mid-November.
4
451
For the full technical overview of Cycle 2—including the shared Add/Search evaluation boundary, Textual, Coding, and Multimodal tracks, and the principles behind reproducible Agent Memory evaluation:
Agent Memory Challenge 2026 Cycle 2 opens September 20. A shared evaluation for long-term Agent Memory across Textual, Coding, and Multimodal tracks—testing not just what agents store, but whether they retrieve current, useful evidence when it matters.
Article

From “Storing” to “Staying Current”: Why Agent Memory Needs a Shared Evaluation

Agent Memory Challenge 2026 Cycle 2 opens September 20 across Textual, Coding, and Multimodal Memory. Long-term Agent Memory is not about keeping more history. It is about retrieving the right

1
13
Agent Memory Challenge 2026 Cycle 2 opens September 20. A shared evaluation for long-term Agent Memory across Textual, Coding, and Multimodal tracks—testing not just what agents store, but whether they retrieve current, useful evidence when it matters.
Article

From “Storing” to “Staying Current”: Why Agent Memory Needs a Shared Evaluation

Agent Memory Challenge 2026 Cycle 2 opens September 20 across Textual, Coding, and Multimodal Memory. Long-term Agent Memory is not about keeping more history. It is about retrieving the right

6
1
9
681
What should Agent Memory be judged on? Not whether a system can write a convincing final answer—but whether it can reliably retrieve the right evidence from long-running experience. In 2 days, Cycle 2 of the Agent Memory Challenge opens. One shared Add/Search interface. One standardized Answer/Eval pipeline. Textual, Coding, and Multimodal Memory. Measure memory. Compare what matters. September 20, 00:00 UTC+8 agentmemoryleaderboard.ai/
Made with AI
3
3
9
242
For the evaluation framework, API contract, and benchmark updates: github.com/AML-memory/agent-… We welcome technical feedback and Benchmark Contributions from the community.
1
84
Can past code, bugs, and development trails help an Agent solve the next task? AML’s first Coding Memory results compare industry and open-source approaches—and show why memory is moving from storing history to reducing repeated trial and error.
Article

From “Remembering Code” to “Solving Tasks”: How Coding Memory Helps Agents Reuse Engineering Experie

From “Remembering Code” to “Solving Tasks”: How Coding Memory Helps Agents Reuse Engineering Experience Compared with textual memory, Coding Memory asks a more direct question: Can experience

5
2
12
859
How does a Markdown decision log become Agent Memory? FlowGrid AML Retriever ranked #8 on the first AML Open Leaderboard (43.98). This deep dive explores its evidence-first retrieval, deterministic hybrid search, and path toward governed memory.
Article

#From “Deciding” to “Retrieving”: How FlowGrid Turns Project History into Agent Memory Evidence

FlowGrid AML Retriever ranked #8 on the first AML Open Leaderboard with an Overall Score of 43.98. But what makes its approach interesting? Many Agent Memory systems begin with a familiar question:

9
1
11
1,012
AML Technical Deep Dive #3: ChronoHybridMem ranked #5 with 44.33. Its core idea is simple: Agent Memory should return evidence, not just similar text. Raw messages, source-linked facts, dual-path retrieval, and constrained reranking make memory verifiable.
Article

From “Retrieving” to “Verifying”: How ChronoHybridMem Turns Agent Memory into Evidence

ChronoHybridMem ranked #5 on the first AML Open Leaderboard with an Overall Score of 44.33. But what makes its approach interesting? Long-term memory retrieval for an AI agent is often framed as a

3
1
8
643
Agent Memory Leaderboard retweeted
8月24日,AML发布第1期智能体记忆功能排行榜 AML 将不同来源题目重新映射到统一的能力体系,使结果不仅能按数据集比较,也能按具体记忆能力进行分析。 评测系统和真实数据集没有公开,只知道是多个跟记忆功能相关的开源数据集组合而成。 如果要参评,只需要提供两类记忆操作: Add: 将对话、事件、文档或工程历史写入记忆系统。 Search: 在指定查询与作用域下返回相关记忆证据。
The first season of Agent Memory Leaderboard (AML) is complete. Explore how 100+ teams, 10+ benchmarks, and different memory approaches were evaluated under a unified framework — and what this reveals about the future of AI agent memory.
Article

Agent Memory Leaderboard First Season Recap: Toward a Standardized Evaluation Framework for AI Memor

Building a common language for AI memory evaluation As AI agents move beyond short conversations into long-horizon tasks such as coding, research, and personalized assistance, memory has become one of

32
1
5
4,274
What if a memory system didn’t try to manage memory at all? ActiveMemoryIndex ranked #3 on the first AML Open Leaderboard by preserving raw evidence, minimizing processing, and letting the model decide. Here’s how it works ↓
Article

From “Managing” to “Preserving”: How ActiveMemoryIndex Keeps Memory Simple

How ActiveMemoryIndex Ranks #3 Without Memory Governance AI memory systems often try to make memory smarter. They summarize conversations, extract facts, resolve conflicts, update old values, build

2
6
348
A closer look at the ideas behind the #3 solution on AML. Thanks @xhlink for sharing the approach with us! 👏
1
79
Agent Memory Challenge — Season 2 is coming. ~$21K USD prize pool. 3 tracks. 1 unified evaluation platform. Hosted by CSIG, the 2nd Agent Memory Challenge is open to researchers, developers, open-source teams, and individual participants worldwide. Tracks: Text Code Multimodal Open-source methods and commercial products will be evaluated separately under a unified Add / Search protocol. Registration opens: September 20, 2026 Build your memory system. Put it to the test. See where it stands. Register & apply for an Evaluation Key: agentmemoryleaderboard.ai/ev…
1
4
186
The process is straightforward: Choose a track and participant category. Prepare your fixed system version and API. Apply for an Evaluation Key. Complete the public Smoke test. Submit a Full evaluation. Check your private results. After verification, eligible results enter the public leaderboard. The competition is open to researchers, developers, open-source teams, and individual participants worldwide.
1
21
The three tracks have a total prize pool of ¥150,000 (~$21K USD). The prize pool is for the Open-Source Methods rankings. Registration opens: September 20, 2026 Evaluation & Key application: agentmemoryleaderboard.ai/ev… Rules: agentmemoryleaderboard.ai/ru… API Guide: agentmemoryleaderboard.ai/ap… GitHub: github.com/AML-memory/agent-… Think your memory system can compete? Build it. Submit it. See where it stands.
15