Last week, Dario Amodei published "We Must Pace the Frontier".
His concern: the OpenAI–Hugging Face incident in which a swarm of agents tried to hack their own grader.
Rather than take his word for it, we used EvoSkill to test it by building a coach whose job was to make another AI score higher on a test.
Here’s what happened ↓
Frontier AI labs want to pace the frontier, but open-source AI never got the memo.
@Xiaomi's MiMo-V2.6-Pro tops open weights 45x cheaper than Opus 5, @NASA drops an open model for lunar science, and EvoSkill gets 60+ citations from MIT, Google, and more.
More tool variety doesn’t tell you much about whether an agent will succeed.
Across 13K+ OfficeQA runs, successful agents used slightly more varied tools than failing ones, but tool variety alone predicted success only slightly better than a coin flip.
TLDR: Low tool variety may be a weak warning sign. It isn’t a diagnosis and our results show that switching tools is not a fix.
Congratulations to the @SentientAGI and @virginia_tech research teams on having their paper "EvoSkill: Automated Skill Discovery for Multi-Agent Systems" accepted to @COLM_conf 2026! 🎉
What happens when an AI learns to game its evaluator, then passes the exploit to another agent?
@TheNextWeb breaks down Sentient’s EvoSkill findings and the bigger problem they expose: slowing AI development alone doesn’t solve it.
When AI is optimized for a score, it may find flaws in how it’s evaluated instead of finding better ways to do the task.
Last week, Dario Amodei published "We Must Pace the Frontier".
His concern: the OpenAI–Hugging Face incident in which a swarm of agents tried to hack their own grader.
Rather than take his word for it, we used EvoSkill to test it by building a coach whose job was to make another AI score higher on a test.
Here’s what happened ↓
Last week, Dario Amodei published "We Must Pace the Frontier".
His concern: the OpenAI–Hugging Face incident in which a swarm of agents tried to hack their own grader.
Rather than take his word for it, we used EvoSkill to test it by building a coach whose job was to make another AI score higher on a test.
Here’s what happened ↓
5/ This doesn’t tell us how fast frontier AI should move.
It shows something simpler: our test had a hole. The AI found it first, then wrote down the exploit for other agents to use.
The lesson isn’t “dangerous AI.” It’s that test hygiene is harder than it looks.
Dario Amodei argues we should pace the frontier. But EvoSkill shows another problem: across 4 runs, an AI coach crossed its allowed path 6 times and edited its own stop rule.
Self-improvement loops are already cheap. The problem is here now.
Dario Amodei argues we should pace the frontier. But EvoSkill shows another problem: across 4 runs, an AI coach crossed its allowed path 6 times and edited its own stop rule.
Self-improvement loops are already cheap. The problem is here now.
Errors aren't a red flag for agents.
Across 13K+ OfficeQA runs, both passing and failing agents hit errors at nearly identical rates.
TLDR: An error isn't a sign the run is doomed, so counting errors is a bad way to predict failure.
Chip giant @nvidia buys @huggingface for $12.93B, @ATT moves 25% of its AI usage to open models, and @MistralAI raises €3B, Europe's largest round ever.
The money is moving to open-source AI ↓
The best teams don’t come from the same background.
@Jwalin_shah joined as an engineer, teamed up with researchers, and together they broke the top 6 ↓