Duncan Curtis retweeted
Meet @DuncanCTech, our SVP of Product and Operations. 👋 Previously, Duncan was Head of Product at @zoox, SVP at @SamaAI, and Product Lead at @Google Play. He was also the creator of Fruit Ninja and Jetpack Joyride. Fun fact: Duncan loves figuring out how physical things work, so when an appliance breaks at home, he usually takes it apart and fixes it himself 🔧. Handyman 2.0.
2
3
18
703
Duncan Curtis retweeted
I don’t often write up my thoughts, but here’s a new blog post summarizing some of my recent musings on RSI and verification: Recursive Self-Inflation. A lot of it is informed by my own research experience and by grappling with the fact that the strongest recent RSI results depend, and quite reasonably so, on fast, automatic feedback: we generate candidates, evaluate them, retain the ones that work, and keep iterating. The problem, however, is that every practical verifier has holes in it, and sustained optimization will eventually find and exploit them. That matters a bit differently in a recursive loop than in a static eval. In a static eval, a verifier mistake just gives you the wrong score. In an RSI loop, an accepted mistake becomes the seed for the next round, so errors can compound over time while the curve still looks like the model is progressing. I recently went through evidence of verifier failure across coding and science benchmarks, reward hacking, model judges, and formal proofs, and tried to characterize what I think credible RSI results would need to report, and what they need to rule out. As always, feedback welcome! mariya.fyi/posts/recursive-s…
6
15
164
8,704
Can you out speedrun AI? Give it a go.
Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games. We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon. On simpler games, agents get close to the human world record. However, on more complicated games like Pokemon Blue, the best models Kimi-K3 and Opus 5 are ~4x off the world record. Benchmark, paper, and demo below.
4
244
Duncan Curtis retweeted
Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games. We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon. On simpler games, agents get close to the human world record. However, on more complicated games like Pokemon Blue, the best models Kimi-K3 and Opus 5 are ~4x off the world record. Benchmark, paper, and demo below.
9
25
134
18,573