The bottleneck in agentic AI just moved again, and this time it's endurance, not intelligence.
For two years the question was which model is smartest. Now it's which agent you can leave alone for a week and trust to still be working when you check back.
A feature flag reading gpt-6-astra-aeon turned up in Codex Desktop on September 3, the day OpenAI shipped GPT-6 Astra at ten and fifty dollars per million tokens. The Agents API went into public beta September 10. DevDay is September 29, three days out. Nobody at OpenAI has said the word Aeon publicly, no product page, no pricing, no benchmark. Just a config string and a lot of people convinced it points to a 24 hour agent that keeps a project moving in the background, remembers yesterday, and wakes itself up when something changes.
That rumor didn't come from nowhere. Meta debuted Muse on September 8 with a customizable avatar built into apps used by hundreds of millions of people, and by late September JPMorgan was telling clients it could become the biggest AI app since ChatGPT.
The whole category is making the same bet. xAI's Grok Bot has been in beta since August 11 at more than 120 dollars a seat, an agent that stays on instead of one you reopen every morning. Manus was about to get folded into Meta, that deal collapsed, and it's now raising 500 million dollars on its own at a 4 billion dollar valuation, still betting a standalone long horizon agent is a real business. Cognition, the company behind Devin, went from 492 million to over 1 billion in annualized revenue in about four months, valuation moving from 25 billion in May to 48 billion since, on one idea: an agent you assign a ticket to and check on later, not one you chat with turn by turn.
The claim going around, unverified but consistent with OpenAI's own math announcements, is an internal run of roughly 10,000 coordinating agents working a Millennium Prize problem for 88 straight hours. Exact number or not, it tells you what the labs think the next unit of scale looks like. Not a bigger model. A longer run.
I watched this shift once before in cloud infrastructure, when the hard part stopped being whether a server could handle load and became whether anyone had built the autoscaling around it. The hard part here is no longer whether a model can plan a three day task. It's whether anyone trusts it enough to stop checking in every hour.
That's the bottleneck nobody's pricing yet: review capacity. An 88 hour agent run produces more decisions than any person can audit live. The product that wins this round won't be the smartest agent. It'll be the one with the clearest record of what it did and why, because that's the only thing that lets a person actually walk away.
nitter.net/CodexResets1/status/21…
🚨 OpenAI Set to Reveal a Long-Term Agent at DevDay
- Codenamed "Aeon" — built for long-running tasks, similar to Grok Bot or Manus, working for hours, days, even weeks
- Likely built on Astra, already strong at long-horizon work
- Runs in a cloud environment like Cursor — sets everything up remotely and keeps grinding until the task's done
- OpenAI already has the infra (hosted sandboxes, multi-agent workflows) to make this the natural next step
- Rumored to support multiple agents collaborating on the same task — if real, that's a big deal