So proud to have helped build ARC-3. Can't wait to see what people throw at it!
Announcing ARC-AGI-3 The only unsaturated agentic intelligence benchmark in the world Humans score 100%, AI <1% This human-AI gap demonstrates we do not yet have AGI Most benchmarks test what models already know, ARC-AGI-3 tests how they learn
6
200
Fraser Scott retweeted
Today we're launching the ARC-AGI-3 Toolkit Your agents can now interact with environments at 2,000 FPS, locally. We're open sourcing the environment engine, 3 human-verified games (AI scores <5%), and human baseline scores. ARC-AGI-3 launches March 25, 2026.
17
65
438
191,383
Don’t live your life pre-AGI like an any% speedrun
2
211
redesigning a game for the 4000th time
1
4
190
Fraser Scott retweeted
Measuring Agents With Interactive Evaluations ARC Prize's OpenAI DevDay 2025 presentation is now live Link below
ARC Prize Foundation + @OpenAI DevDay 2025 @arcprize was excited to present as the only external benchmark at OpenAI's DevDay @GregKamradt presented "Measuring Agents With Interactive Evaluations" He covered why interactive benchmarks give us new tools to measure intelligence
2
2
57
14,652
Fraser Scott retweeted
ARC-AGI-3 Preview: +3 Games Released We’ve opened 3 previously private holdout games from the Preview Agent Competition Now 6 games are available to play online and via Agents API Each game was selected to expand the novelty of ARC-AGI-3 public games Can you beat them?
11
79
341
44,915
This fits my experience. Models output a flowery CoT that sounds like something eg. a chess grandmaster would say, but omits basic reasoning steps. For example, GPT-5 *knew the rules* of this puzzle, but failed to properly apply them or even self-correct three.arcprize.org/replay/ls…
Is Chain-of-Thought Reasoning of LLMs a Mirage? ... Our results reveal that CoT reasoning is a brittle mirage that vanishes when it is pushed beyond training distributions. This work offers a deeper understanding of why and when CoT reasoning fails, emphasizing the ongoing challenge of achieving genuine and generalizable reasoning. ... Our findings reveal that CoT reasoning works effectively when applied to in-distribution or near in-distribution data but becomes fragile and prone to failure even under moderate distribution shifts. In some cases, LLMs generate fluent yet logically inconsistent reasoning steps. The results suggest that what appears to be structured reasoning can be a mirage, emerging from memorized or interpolated patterns in the training data rather than logical inference. ... Together, these findings suggest that LLMs are not principled reasoners but rather sophisticated simulators of reasoning-like text.
1
2
447
Fraser Scott retweeted
View other GPT-5 ARC-AGI-3 replays * vc33: three.arcprize.org/replay/vc… * ft09: three.arcprize.org/replay/ft… * ls20 Guided (GPT-5 was told what the rules of the game were): three.arcprize.org/replay/ls…
1
2
31
9,206
I don't see how Genie 3-style world models will ever be interesting It is simply impossible to simulate our universe in this open-form way; if you can interact with the simulation, you can push it into a regime where the model's understanding breaks down
2
166
and if it never breaks down, congratulations, our universe is holographic and you have created a computer that can simulate the universe from inside and predict the future
58
a random agent plays @arcprize ls20
1
133
The program to beat a given ARC game is surprisingly short.. if you are intelligent enough to find it
Many of the SOTA solutions on ARC-AGI-1/2 were advanced by ARC-AGI domain specific languages (DSLs) We knew these would come around for ARC-AGI-3 but not this quick Someone already created one to beat ft09 We don't expect this to generalize to private games, but interesting to see how competition dynamics play out
1
141
interpreter is the only job that has any chance of being automated under the current paradigm and even then you’d still want someone around to blame any cultural foopahs on
Microsoft just dropped a study showing the 40 jobs most at risk by AI and the 40 most secure.
1
181
The Online Safety Act is the biggest assault on free speech in modern British history. Unless we resist it, Britain will become a digital surveillance state. Section 179 of the act makes it a criminal offence to say something false that causes 'non-trivial psychological harm,' criminalising humour, satire, political dissent and 'offensive' tweets. Section 44 gives a single government minister the power to rewrite the censorship rules without Parliament. This means that a single person can direct Ofcom to impose new censorship obligations on the internet, and platforms will be legally forced to comply. Britain is sleepwalking into a digital dictatorship. This in turn paves the way for a social credit scoring system in which counter narrative views are crushed, and only government approved views are allowed. Join the hundreds of thousands of Brits who are fighting back. You can help by signing the petition here 👇 petition.parliament.uk/petit…
185
2,091
6,456
361,246
wen AGI?
11% 2025
22% 2026
22% 2027
44% 2028+
9 votes • Final results
92
Some news: We're building the next big thing — the first-ever AI-only social video app, built on a highly expressive human video model. Over the past few weeks, we’ve been testing it in private beta. Now, we’re opening early access: download the iOS app to join the waitlist, or skip the line with an invite code. Codes are limited — if you know someone with early access, they might have one for you. apps.apple.com/gb/app/pika-s…
1
2
300
input data: 64x64 array of pixels output data: an action loss function: score
Every time someone proposes training a model for some domain (virtual cells/materials/whatever), the first sentence should be: input data: x, output data: y, loss function: z. Would make it a lot easier for dum dums like me to understand what exactly is being proposed.
1
188