To ensure that Artificial General Intelligence is open-source and not controlled by any single entity. @SentientEco @OpenAGISummit

San Francisco, CA
Based in United Arab Emirates
Last week, Dario Amodei published "We Must Pace the Frontier". His concern: the OpenAI–Hugging Face incident in which a swarm of agents tried to hack their own grader. Rather than take his word for it, we used EvoSkill to test it by building a coach whose job was to make another AI score higher on a test. Here’s what happened ↓
20
14
68
48,135
BREAKING: OPEN WEIGHT MODELS RAN 56% OF ALL TOKENS ON AI GATEWAY, TAKING THE LIONSHARE OF AI VOLUME.
Frontier AI labs want to pace the frontier, but open-source AI never got the memo. @Xiaomi's MiMo-V2.6-Pro tops open weights 45x cheaper than Opus 5, @NASA drops an open model for lunar science, and EvoSkill gets 60+ citations from MIT, Google, and more.
Article

This Week In OS AI Field Notes: 25 September 2026

Open-Source Models NASA-IBM MODEL: @NASA and @IBMResearch released the Lunar Foundation Model, among the first open models built for lunar science, trained on roughly 2 million image tiles drawn

10
4
46
9,120
A bigger toolbox is not the solution ↓
More tool variety doesn’t tell you much about whether an agent will succeed. Across 13K+ OfficeQA runs, successful agents used slightly more varied tools than failing ones, but tool variety alone predicted success only slightly better than a coin flip. TLDR: Low tool variety may be a weak warning sign. It isn’t a diagnosis and our results show that switching tools is not a fix.
10
1
31
10,167
Congratulations to the @SentientAGI and @virginia_tech research teams on having their paper "EvoSkill: Automated Skill Discovery for Multi-Agent Systems" accepted to @COLM_conf 2026! 🎉
22
10
78
14,153
Explore all of our papers and where they've been published: sentient.xyz/research
8
489
What happens when an AI learns to game its evaluator, then passes the exploit to another agent? @TheNextWeb breaks down Sentient’s EvoSkill findings and the bigger problem they expose: slowing AI development alone doesn’t solve it. When AI is optimized for a score, it may find flaws in how it’s evaluated instead of finding better ways to do the task.
Last week, Dario Amodei published "We Must Pace the Frontier". His concern: the OpenAI–Hugging Face incident in which a swarm of agents tried to hack their own grader. Rather than take his word for it, we used EvoSkill to test it by building a coach whose job was to make another AI score higher on a test. Here’s what happened ↓
14
4
48
11,619
Last week, Dario Amodei published "We Must Pace the Frontier". His concern: the OpenAI–Hugging Face incident in which a swarm of agents tried to hack their own grader. Rather than take his word for it, we used EvoSkill to test it by building a coach whose job was to make another AI score higher on a test. Here’s what happened ↓
20
14
68
48,135
5/ This doesn’t tell us how fast frontier AI should move. It shows something simpler: our test had a hole. The AI found it first, then wrote down the exploit for other agents to use. The lesson isn’t “dangerous AI.” It’s that test hygiene is harder than it looks.
1
10
538
Read the full article on X ↓
Dario Amodei argues we should pace the frontier. But EvoSkill shows another problem: across 4 runs, an AI coach crossed its allowed path 6 times and edited its own stop rule. Self-improvement loops are already cheap. The problem is here now.
Article

Dario Amodei warned about AI gaming its grader. Here's what happened when ours did.

On September 12 2026, Dario Amodei published "We Must Pace the Frontier" and among his concerns is the OpenAI–Hugging Face incident, in which a swarm of agents, among other things, tried to hack the

10
717
Dario Amodei argues we should pace the frontier. But EvoSkill shows another problem: across 4 runs, an AI coach crossed its allowed path 6 times and edited its own stop rule. Self-improvement loops are already cheap. The problem is here now.
Article

Dario Amodei warned about AI gaming its grader. Here's what happened when ours did.

On September 12 2026, Dario Amodei published "We Must Pace the Frontier" and among his concerns is the OpenAI–Hugging Face incident, in which a swarm of agents, among other things, tried to hack the

23
5
56
13,543
Hidden in plain sight 🎨
There's a signature hidden in this painting. Not in the corner, not on the edges, but in the portrait itself. Can you find it?
21
6
62
11,984
Errors don't mean your agent is failing. That's just how they work ↓
Errors aren't a red flag for agents. Across 13K+ OfficeQA runs, both passing and failing agents hit errors at nearly identical rates. TLDR: An error isn't a sign the run is doomed, so counting errors is a bad way to predict failure.
12
1
40
12,522
Money talks 💸 And this week it said open-source AI ↓
14
3
44
12,114
Compute ≠ Provenance
Compute buys capability. But it doesn't tell anyone where the model came from.
13
4
43
12,596
What do you get when an engineer and two researchers team up at an open source AI hackathon? Top 6 in the Arena ↓
The best teams don’t come from the same background. @Jwalin_shah joined as an engineer, teamed up with researchers, and together they broke the top 6 ↓
13
3
32
11,266