Raven 0.2.0 — The Harness of Harnesses, built for RSI. 🐦‍⬛ One harness can't be best at everything. Raven combines its own specialist harnesses (Research, Code, Design, Oncall) with the agents you already use (Claude Code, Codex and more) into one team. And it's built for RSI, and not just at the skill level. The whole harness can be rewritten by AI: prompts, policies, strategy code, playbooks. Every sub-harness, including the orchestration layer itself, is its own instance that can be improved. With Raven you can: 1. Orchestrate many agents as one team. Raven's sub-harnesses and external agents work in one task graph with shared memory across sub-agents, powered by leading orchestration (0.963 Node F1 on the Multi-Agent Orchestration Benchmark). 2. Run long, complex tasks. Oncall and proactive execution keep work going for days, from scientific research loops to shipping a full Godot game. 3. Build vertical agents with RSI. Use Raven's RSI to develop and refine an agent for your domain, and we'll optimize it with you. Experimental for now; reach out to the Raven team(Discord:discord.gg/VVSm8MESch). More in the video and slides below. Open source, Apache-2.0. github.com/EverMind-AI/Raven (lots of work made with Raven lives there, and much of this launch's material was made with Raven too)
598
333
759
281,822
邓亚峰 retweeted
EverOS just crossed 12,000 GitHub stars. Usage has been accelerating too: • 374 downloads yesterday • 2,154 in the past week • 5,926 in the past month Since v1.0.0, we’ve been focused on making EverOS easier to try and more dependable to run. • everos demo now launches a full-screen memory walkthrough with no API key or server setup. • One command lets you watch a memory move through ingest → extract → index → recall. • When you’re ready to build, one OpenRouter key is enough to start the core memory flow: durable Markdown storage, cascade indexing, and keyword search. • Recent releases have also added /api/v2, OpenTelemetry tracing, safer agent-skill extraction, and more reliable background maintenance. Thank you to everyone who tried EverOS, opened an issue, contributed code, or built something with it. More coming soon. github.com/EverMind-AI/EverO…
3
17
8,003
邓亚峰 retweeted
彩蛋一下:Wiki 和 Dreaming 很快就会上线。 这周末在新加坡跟一圈做出海 Agent 的朋友深聊下来,有个很直观的感觉:大家对 Memory 的需求已经从「可有可无」变成了刚需。 我们一开始做 Memory,其实还是比较工程师视角的: 一方面,它可以替代一部分 RAG,不是什么东西都要塞进知识库再检索一遍; 另一方面,它能帮模型省 Token,不用每次把所有历史再喂一遍。 但今年再聊,发现大家关心的问题已经变了。 现在真正重要的,不是「怎么让 Agent 记住更多东西」,而是:怎么让一个 Agent 用久了之后,真的越来越懂你。 比如,它要知道你习惯怎么写东西,知道你做产品时在意什么,知道哪些信息和设定已经发生了变化,知道哪些坑上次踩过、下次应该自动绕开。 这也是为什么,最近「妈生感」这种说法会这么流行。 走到这一步,Memory 就不再只是个存储层,而会变成 Agent 的一套自我更新机制。 这里面我个人觉得最有意思的是 Dreaming,也是我们马上会上线的一个特性,我们内部叫 Reflection。 它不是让 Agent 在那里「玄学做梦」,而是在每次任务结束之后,系统性回顾这次发生了什么:哪些信息值得长期留下,哪些偏好被进一步确认了,哪些状态已经过期失效了,哪些失败应该被总结成下次的经验。 如果没有这一层,Agent 每次都像是「重新开机」;有了这一层,它才会慢慢长出连续性。 接下来,EverMind 会重点做两件事: 一是 Knowledge Wiki,把 Memory 变成用户看得见、改得了、用得上的知识库; 二是 Dreaming,在后台持续做 Reflection,把零散的对话和任务经历,沉淀成更稳定的 Profile。 越往下做,我越确信:Memory 最重要的,不是帮 AI 记住更多,而是让一个 Agent 有机会,越用越像「你的 Agent」。
周末花了很多时间更新 EverOS 的 Repo,算是达到自己想要的状态了。 EverOS 1.0.0 这次不是单点功能更新,是把 agent long-term memory 往“可运行、可审计、可扩展”的基础设施方向推进了一步。 这次最关键的 6 个特性:
3
4
21
3,558
邓亚峰 retweeted
The amazing work of our new designer and developers.
EverOS passed 6,000 stars on GitHub today. This milestone means a lot to our team. We’ve put serious work into building an open home for long-term and self-evolving memory in agents: real use cases, runnable methods, benchmarks, and tools developers can build on. If you believe agents need memory to become truly useful, we’d love your support.
1
2
109
邓亚峰 retweeted
我们新设计了 EverOS 的界面, 现在主要想突出两个点。 第一,它是面向 AI Agent 的“长程记忆操作系统”。 传统 LLM 应用更多依赖短期上下文,而 EverOS 试图把 agent 过去的会话、用户偏好、项目资料和全局知识沉淀成持续可检索的记忆层。文本、PDF、图片、网页等多模态数据可以统一存储,再通过向量、关键词和多模态对齐的混合检索做 mRAG。 目标是让 agent 不再每次从零开始,而是能跨很长时间理解用户、项目和知识背景。 第二,也是我最关注的,是 Skill Self-Evolution。 EverOS 会把过去任务里的轨迹自动抽取成案例,再通过语义聚类形成可复用的技能 SOP,并在后续任务中持续迭代这些技能。 在 EverMind 自建的 EvoAgentBench 评测里,这套自进化机制让一个 27B 模型在复杂软件工程任务上的成功率相对提升约 234.8%,在该评测设置下接近 397B 模型的表现。 这背后的核心判断是:agent 的能力提升,不一定只靠更大的模型,也可以来自长期记忆、经验复用和技能沉淀。
6
5
39
6,789
The more you use it, the smarter it gets — it evolves on its own.
Today we're launching everme — personal memory for the age of agents. The idea: you'll use many different agents, but your memory should belong to you. Just give your agent one sentence, and everme manages your memory across all of them — with self-evolving skills, zero maintenance, and a 2-minute cold start. With everme: • Claude Code and Codex share memory • openclaw and hermes agent share memory • memory travels across platforms and machines → everme.evermind.ai
1
11
449
Today we're launching everme — personal memory for the age of agents. The idea: you'll use many different agents, but your memory should belong to you. Just give your agent one sentence, and everme manages your memory across all of them — with self-evolving skills, zero maintenance, and a 2-minute cold start. With everme: • Claude Code and Codex share memory • openclaw and hermes agent share memory • memory travels across platforms and machines → everme.evermind.ai
1
8
638
To evaluate your Claw agent's evolutionary capabilities, you can utilize EvoAgentBench, the second benchmark dedicated to agent evaluation on Hugging Face.
~200 downloads in under a week. Blown away. Thank you.
5
501
EverMind is Hiring: Technical PM (Agent OS & Memory) Based in Silicon Valley | Shanghai | Beijing What we need: Tech + Product: Deeply understand Agent execution mechanics under the hood (OpenClaw/Hermes Agent is a +). AI-Native: Fluent in Vibe Coding. You can spin up working demos yourself to validate concepts. LTM Focus: Intensely driven to build an Agent OS with true Long-Term Memory. Flexible setup: Open to part-time/intern as a trial. DM me to chat!
15
11
46
35,988
AI self-evolution is undoubtedly the ultimate core of next-gen AI—and true evolution is fundamentally built on Long-Term Memory. 🧠 At EverMind, we believe every AI of the future must have long-term memory. If you want to build an AI that continuously adapts and unlocks a massive data flywheel, you need EverOS. ♾️ The update we’ve been polishing for ages is FINALLY live! 🚀 It's way more than just a product update—Methods, Benchmarks, Usecases, and a fresh site are all here. Here’s the breakdown: 👇 1️⃣ EverMemOS ➡️ EverOS: The ultimate one-stop shop. EverOS now empowers your Agents with self-evolution capabilities—just like Hermes Agent. Plus, we’ve added full multi-modal support. Use our Methods to customize Usecases into your own Agents, then benchmark to optimize them. It's the absolute all-in-one king. 👑 2️⃣ EvoAgentBench is LIVE & Open-Source: The perfect tool to test your custom Claude Code, OpenClaw, Hermes, or any Agent you throw at it. 📊 3️⃣ Brand New Website: Aesthetics matter. The new vibe, colors, and interactive experiences are absolutely off the charts. 🎨🔥 Dive in here: github.com/EverMind-AI/EverO… #EverOS #AI #AgentTools #Harness @Memory #OpenClaw #Hermes #ClaudeCode
1
419
Every major wave of computing has been defined by how we store and retrieve information. Mainframes, databases, the cloud. AI is no different. The teams that competed at Memory Genesis 2026 understand something most of the industry has not fully internalized yet: memory is not just infrastructure. It is intelligence itself. This event was a glimpse of that future. Grateful to everyone who showed up to build it with us.
The Memory Genesis Competition 2026 Final Event kicked off today in Mountain View, CA. Hosted by @shanda_group and @evermind, and supported by OpenAI and AWS, the Memory Genesis Competition 2026 brought together innovators, researchers, investors, and builders to explore how next-generation memory technologies will define the future of AI. A landmark day for the memory industry. Here is what went down. (Stay until the end. You will want to see how this room looked.) 🧵
2
429
邓亚峰 retweeted
A few weeks ago we published our Memory Sparse Attention paper, a new way to give AI models long-term memory that actually works. Today's LLMs/Agents forget. They can only hold so much context before things start falling apart. We built a system that lets a model remember up to 100 million tokens, the length of about a thousand books, and still find the right answer with less than 9% performance loss. On several benchmarks, our 4-billion parameter model even beats RAG systems built on models 58× its size. The idea? Instead of searching a separate database and hoping the right info comes back (that's how RAG works), we built the memory directly into how the model thinks. It learns what to remember and what to ignore, end to end, no separate retrieval pipeline needed. The response to the paper blew us away. Researchers and engineers everywhere asking the same thing: "When can we see the code?" So we got to work, cleaned up the inference code, documented it, and made it ready for the community to dig in. You asked for it. We open-sourced it. github.com/EverMind-AI/MSA
6
21
125
13,883
邓亚峰 retweeted
稍微剧透一下 1. @EverMind 的 MSA 论文被 AlphaXiv 选中发表了,流量还非常不错 2. MSA 的 Inference 本周会开源 github.com/EverMind-AI/MSA
Scaling Attention to 100M context!? Memory Sparse Attention introduces an idea where instead of rereading an entire 100M-token entry, it learns to jump straight into the relevant memories and reason from them end-to-end. More specifically, it first encodes documents into compressed memory slots, then for each question it uses a learned router to score which chunks are actually relevant, pulls only the top few, and runs normal attention over that tiny assembled context. So the model’s compute grows with “how much it needs to look at” not “how much memory exists”. This retrieval step is trained jointly with answer generation, so memory lookup is part of the model itself, and can decouple memory capacity from reasoning cost.
7
10
80
12,829
MSA这么基础的工作,还得到了大家这么高的关注。 GitHub 链接(欢迎继续点星⭐️):github.com/EverMind-AI/MSA 论文链接:zenodo.org/records/19103670
1
6
289
MSA (Memory Sparse Attention) represents our significant exploration in the field of long-term memory. It stands as the first end-to-end long-term memory framework for large models to genuinely achieve a 100M context length. Interestingly, as the memory length scales from 16K to 100M, the model's performance score decreases by a mere 9%, demonstrating highly robust scalability. Main contribution: 1,We propose MSA, an end-to-end trainable, scalable sparse attention architecture with a document-wise RoPE that extends intrinsic LLM memory while preserving representational alignment. It achieves near-linear inference cost and exhibits < 9% degradation even when scaling from 16K to 100M tokens. 2,We introduce KV cache compression to reduce memory footprint and latency while maintaining retrieval fidelity at scale. Paired with Memory Parallel, it enables high-throughput processing for 100M tokens under practical deployment constraints, such as a single 2×A800 GPU node. 3,We present Memory Interleave, an adaptive mechanism that facilitates complex multi-hop reasoning. By iteratively synchronizing and integrating KV cache across scattered context segments, MSA preserves cross-document dependencies and enables robust long-range evidence integration. 4,Comprehensive evaluations on long-context QA and Needle-In-A-Haystack benchmarks demonstrate that MSA significantly outperforms frontier LLMs, state-of-the-art RAG systems and leading memory agents. Welcome to feedback: github.com/EverMind-AI/MSA zenodo.org/records/19103670 We are looking for passionate talents to join our team! If you are interested in our work and vision, please don't hesitate to send us an email at evermind@shanda.com.
2
1
28
2,960