Simulation research and infrastructure for human-aligned AGI patronus.ai

1/ Today, we are thrilled to announce Generative Simulators, a new class of adaptive, auto-scaling environments for AGI training and evaluation 🤖🧵 Static datasets, hand-authored environments, and human-curated demonstrations do not automatically scale with the learning patterns of the trained model. We propose Generative Simulators as a principled alternative: environments that evolve, evaluate, and adapt to agent behavior over time. Technical Report: patronus.ai/generative-simul… Blog: patronus.ai/blog/introducing…
7
24
71
19,612
We're excited to share that our paper on prompt infilling for dLMs has been accepted to the "Non-Autoregressive Language Models for Fast & Flexible Text Generation" workshop at COLM 2026 🎉 Masked diffusion language models like LLaDA and Dream have bidirectional architectures, yet they can't infer prompts from desired outputs out of the box. We traced this to a training gap: during SFT, prompts are simply never masked. The fix is straightforward. Mask both prompts and responses during SFT (full-sequence masking). That single change unlocks the ability to infer task-adapted prompts from few-shot examples. Prompt infilling in dLMs just requires the right training approach. We're calling on the community to release full-sequence masked SFT checkpoints alongside standard ones. Congrats to @akkikiki and Keisuke Sakaguchi! Paper: arxiv.org/abs/2604.03677 Fine-tuned models: huggingface.co/collections/a…
4
4
35
1,689
Meet @DuncanCTech, our SVP of Product and Operations. 👋 Previously, Duncan was Head of Product at @zoox, SVP at @SamaAI, and Product Lead at @Google Play. He was also the creator of Fruit Ninja and Jetpack Joyride. Fun fact: Duncan loves figuring out how physical things work, so when an appliance breaks at home, he usually takes it apart and fixes it himself 🔧. Handyman 2.0.
2
3
18
692
We're hosting a researcher night next Wednesday at our office in SF. Fast-paced spotlight talks on post-training and evaluation. ⚡️ The talks: @mariyaivasileva on Beliefload: an evaluation of how LLMs formulate and revise hypotheses in light of new evidence, and how they diagnose and repair a broken environment without overwriting the parts that already work. @Nick_saban20 on GlobeBench: a benchmark for whether language models can faithfully simulate the environments we train agents in. @AnmolGulati06 on Beyond Rows to Reasoning: an agentic framework for reasoning over and editing enterprise spreadsheets with millions of cells, cross-sheet dependencies, and embedded charts. @akkikiki on SpeedrunBench: a benchmark that asks not whether an agent can finish a game, but how fast. Drinks, sushi, merch incoming.
1
6
21
2,039
PatronusAI retweeted
True story: our original company name when we were in stealth was Zeno AI, not Patronus AI. Back in 2023, Rebecca and I started the company on the idea that AI evaluation was like Zeno’s Paradox. Every time models reached the bar we had set, the frontier of what we needed to evaluate moved further ahead. Fast forward to 2026, I still think Zeno's Paradox is the most accurate way to describe AI progress. Every finish line we draw for AI moves the moment it reaches it. In the original Zeno's Paradox, every time Achilles reaches where the tortoise stood, the tortoise has moved a little further on, so he appears to chase forever. Mathematically the paradox does resolve. The shrinking gaps sum to a finite length, and Achilles does pass. In the AI version of the paradox, the idea is that humans are the tortoise, and AI is Achilles. We start far ahead, carrying millions of years of evolutionary and cultural learning. And over the last few years, Achilles has closed most of the visible distance. But the reason the AI version is more interesting is that the tortoise keeps speeding up. Each time AI closes a gap, it hands the people defining the frontier a better tool for defining the next one. Mathematicians use the newest models to probe harder open problems. Researchers use them to help design the very evals the next generation of models will struggle against. Everyone is anxious about the rate of progress in AI right now, but ultimately I think we’re going to land in a pretty optimistic place: AI will enable humans to reach farther than ever. @danshipper has a nice GIF of this tortoise/Achilles setup, see below.
3
5
40
1,556
Meet @chaaelaes from our data quality team. 👋 Chelsea was previously at UC Irvine where studied biomedical engineering. She spent time in a neuroscience lab building imaging hardware and researching neurodegenerative diseases. Fun fact: Chelsea has played the violin since she was 3 years old🎶
2
1
19
1,411
Meet Abdelrahman Madkour, from our research team. 👋 Madkour was a previously a Research Scientist at Meta, where he worked on coding benchmarks for coding agents. Before that, he completed his Ph.D. in CS at Northeastern University, where he focused on procedural content generation in games. Fun fact: Madkour plays the Arabic oud 🪕
1
2
16
692
Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games. We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon. On simpler games, agents get close to the human world record. However, on more complicated games like Pokemon Blue, the best models Kimi-K3 and Opus 5 are ~4x off the world record. Benchmark, paper, and demo below.
9
25
134
18,498
Meet Zhe Li, from our research team. 👋 Previously, Zhe lead post-training at Inflection AI (now Microsoft). Before that, Zhe was a key contributor to FaceID and Apple Vision Pro at Apple. Zhe did his PhD in computer science at the University of Iowa, where he worked on neural network optimization. Fun fact: Zhe is a naturally curious person and will happily go down a rabbit hole researching the most random things.
3
3
43
5,603
Meet @KaranSamel, from our research team. 👋 Karan completed his PhD in machine learning at Georgia Tech in 2024, where his research explored ways to integrate knowledge and structured reasoning processes into AI models to make them more efficient and purpose-driven. During his PhD, he also worked in industry research labs at Google Research, Amazon, and IBM Research. Fun fact: Karan is an avid nature photographer. He once braved an icy mountain drive at 2 a.m. in Iceland just to capture the northern lights. 🏔️
3
4
23
942
Today we're excited to announce FigmaTrace, the first open source computer-use dataset for design. We captured 200+ hours and 3400+ trajectories of real design work in Figma, recording step by step how designers execute complex, long-horizon tasks. Dataset, paper, and blog below.
8
16
92
4,378
We observe that training on FigmaTrace consistently improves step-level action performance by up to 26% on out-of-domain computer-use benchmarks across @AIatMeta Glimmer, @googlegemma Gemma-4, and @QwenDevs Qwen families. We attribute these results to our design-phase based video-to-trajectory conversion method and diverse open and close ended tasks.
1
10
833
Meet @nick_saban20, from our applied research team 👋 Nick joins us from UC Berkeley, where he graduated at 19! His academic research covered safety, alignment, and multimodal systems. Last year, Nick presented his paper AutoAdv at NeurIPS, a multi-turn automated red teaming framework for jailbreaking frontier LLMs. Fun fact: when he's not fine-tuning models, he's fine-tuning cars 🚗
2
2
17
1,052