CS prof @Penn & founder @Rabdos_AI: creating data for the world’s hardest reasoning, from the chalkboard to the wet lab

Philadelphia, PA
Pinned Tweet
Friends, followers, and strangers on X, I recently got excited about mapping the jagged frontier of AI models in math and other STEM areas. With my colleague Prof. Rob Ghrist @prof_g , I founded Rabdos AI rabdos.ai to create original, research-level problems at scale to advance frontier AI capabilities. In the coming weeks, we'll be sharing exciting technical updates through @Rabdos_AI. Please give us a follow if you're interested in model evaluation, AI for math/science, autoformalization in Lean, robustness of world models, and much more.
1
6
34
2,142
Mayur Naik retweeted
The status quo for R1 institutions cannot last. I'm guessing there will be a barbell distribution of focus that splits the classical research/teaching fusion we currently enjoy. A few OOM of improvement in AI and we might see: 1. The research side goes industrial -- applied mathematicians producing enormous bodies of interdisciplinary work; physical science labs running in warehouses with heavy robotics and automated data collection and processing; economics and social science simulations at extreme scales... 2. Teaching is top-quality, high-touch, personalized-service, with courses becoming works-of-art and top professors being superhuman performers and guides. In short, each end of the research-teaching axis taken to extremes. There are hints of this divergence already. At Penn, the two top schools (by $$$ & prestige) are the medical school (no undergraduate teaching; lots of heavy lab-based research at scale), and the business school (almost all teaching, highly profitable, most research occurring in pockets like the stats department at wharton). Ramp that split with AI improving 100-fold and more, and we will see a vast bimodal transformation that integrates the university more with industry (large scale research) and serves students better (extremely attentive teaching).
4
8
76
8,867
Congrats to Anthropic on shipping Claude Opus 5.5! Here's an especially noteworthy article by Boris Cherny on using this model to formally verify the Claude Agent SDK using Lean. Rabdos is an industry leader in creating high-quality Lean data; check out our benchmark at rabdos.ai/math-index/lean
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
1
11
1,008
Mayur Naik retweeted
Congratulations to Anthropic on the release of Claude Opus 5.5! 🎉 At Rabdos, we spend our days thinking about the hardest problems in math and science, so it's exciting to watch frontier reasoning keep advancing this quickly.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
1
2
9
1,267
📣 We're growing our team at Rabdos AI! We're looking for a 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿 to lead our work on mathematical formalization in Lean. You'd turn informal math into formal proofs and build autoformalization pipelines alongside frontier AI models, working with some of the sharpest mathematicians and computer scientists I know. If you love Lean, rigorous math, and hard open problems, I'd love to hear from you. Please share with anyone who comes to mind! 👇
📣 𝗪𝗲'𝗿𝗲 𝗵𝗶𝗿𝗶𝗻𝗴! 𝗥𝗲𝘀𝗲𝗮𝗿𝗰𝗵 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿 — 𝗠𝗮𝘁𝗵𝗲𝗺𝗮𝘁𝗶𝗰𝗮𝗹 𝗙𝗼𝗿𝗺𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻 𝗶𝗻 𝗟𝗲𝗮𝗻 At Rabdos AI, we build research-level STEM tasks that stay beyond the reach of frontier AI models, even as those models approach millennium-prize-level problems. Founded by academics at the University of Pennsylvania, we've bootstrapped to $𝟮𝟱𝗠 𝗶𝗻 𝗔𝗥𝗥 𝗶𝗻 𝗹𝗲𝘀𝘀 𝘁𝗵𝗮𝗻 𝗮 𝘆𝗲𝗮𝗿, and we're growing the team. 𝗪𝗵𝗮𝘁 𝘆𝗼𝘂'𝗹𝗹 𝗱𝗼 🔹 Lead the design and evaluation of methods that turn informal mathematical reasoning into formal proofs in Lean 🔹 Build 𝗮𝘂𝘁𝗼𝗳𝗼𝗿𝗺𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻 workflows and integrate them with frontier AI models 🔹 Curate high-quality formal reasoning datasets alongside mathematicians and computer scientists 🔹 Prototype, experiment, and iterate on tools that make formal verification scalable and reliable 𝗪𝗵𝗮𝘁 𝘄𝗲'𝗿𝗲 𝗹𝗼𝗼𝗸𝗶𝗻𝗴 𝗳𝗼𝗿 ✅ Experience with formal theorem proving and autoformalization, ideally in 𝗟𝗲𝗮𝗻 𝟰 ✅ 𝗚𝗿𝗮𝗱𝘂𝗮𝘁𝗲-𝗹𝗲𝘃𝗲𝗹 knowledge of mathematics or physics ✅ Experience building and improving AI systems such as LLM pipelines or evaluation tools ✅ A degree in CS, Math, Physics, or a related field (or equivalent research experience) ✅ Clear communication across disciplines and a commitment to rigorous validation 𝗕𝗼𝗻𝘂𝘀 𝗽𝗼𝗶𝗻𝘁𝘀 ➕ Strong Python skills ➕ Familiarity with how LLMs are trained, fine-tuned, or evaluated 𝗧𝗵𝗲 𝗱𝗲𝘁𝗮𝗶𝗹𝘀 📍 𝗢𝗻-𝘀𝗶𝘁𝗲 𝗶𝗻 𝗣𝗵𝗶𝗹𝗮𝗱𝗲𝗹𝗽𝗵𝗶𝗮 | 𝗙𝘂𝗹𝗹-𝘁𝗶𝗺𝗲 | Start ASAP 💰 $𝟭𝟱𝟬𝗞 – $𝟮𝟮𝟱𝗞 + 𝗲𝗾𝘂𝗶𝘁𝘆 + performance bonuses paid multiple times a year 🩺 Health insurance, a daily Uber Eats stipend, and a commute benefit If your background is a good fit, we'll reach out within 𝗼𝗻𝗲 𝗯𝘂𝘀𝗶𝗻𝗲𝘀𝘀 𝗱𝗮𝘆 to schedule an intro call. 👉 𝗔𝗽𝗽𝗹𝘆 𝗵𝗲𝗿𝗲: rabdos.ai/apply/research-eng… Know someone who'd be a great fit? Please tag them or share this post! 🙏 Rabdos AI is an equal opportunity employer.
3
2
10
1,269
Are endpoint protection systems like CrowdStrike prepared for large-scale agentic workloads? My research group has been running 100's of codex agents in parallel on a machine this summer. The agents are not coordinating (or at least aren't supposed to): each works on an independent research task in a STEM domain such as math, SWE, materials, biochem, ... In a matter of days, they have started tripping conventional security software in ways that have alarmed IT folks and are very hard to debug. The only solution seems to be to isolate the machine from the network, but that cuts off the agents from even basic web search. This and recent incidents of agents running amok at frontier AI labs present a massive opportunity for research and startups in AI x security. How do we specify / enforce security policies for agents? How do we triage alarms? Seems especially difficult since good vs. bad behavior is superficially indistinguishable : agents exploring how to use a low-level systems tool could be misconstrued as trying to breach security; a sequence of such seemingly random tool calls would pose an even greater challenge. Feels like the early days of the wild west, and a harbinger of things to come ...
2
263
Every year I take an international trip, I measure in-flight internet connectivity as an indicator of global technological progress. 2026 represents a *huge* step up in this measure, entirely single-handedly due to @Starlink. They have essentially solved this problem from 0 to 1.
3
13
1,567
Check out Rabdos Math Index! A holistic benchmark for math x AI that draws from our experience supplying math data to frontier AI labs over the last six months. Comprises graduate-level theorems formalized in Lean and checked by semantic rubrics, research-level problems with single numeric answer, and problems requiring visual reasoning over mathematical figures. DM me or email contact@rabdos.ai for more info!
We are delighted to introduce Rabdos Math Index 📐 One holistic measurement for mathematical reasoning on proof, solution, and perception. Claude Opus 5 leads at 46. GPT-5.6 Sol & Claude Fable 5 follows at 39. No other model scored above 25. Mathematics remains open.
1
12
1,516
Mayur Naik retweeted
THE SYMPOSIUM PUZZLE: The final dinner of the symposium was less a banquet than a convergence theorem that had failed to be uniform. Five luminaries -- Hardy, Poincaré, von Neumann, Gödel, and Ramanujan -- sat in a row at the head table, each in a different jacket, each with a different drink, each newly returned from a different lecture tour, and each guarding a different mathematical instrument as though it were a proof of the Riemann Hypothesis. Hardy sat brooding at the far left in herringbone, one hand curled around an espresso, the other resting upon an antique abacus whose beads he refused, on principle, to move. Immediately to his right sat a severe scholar in charcoal, upright as a metronome and no more companionable. Poincaré, ever the classicist, wore tweed. Farther down the line, Ramanujan (newly back from Göttingen) sat resplendent in navy, sipping tea and turning a golden compass over in his fingers as though it might draw identities straight out of the air. The navy jacket sat immediately to the left of the pinstripes, a juxtaposition that pleased no tailor present. The guest who had lectured at Cambridge, meanwhile, was the one in herringbone. When the conversation turned from foundations to apparatus, the scholar fresh from Princeton began boasting of a brass astrolabe he had recently acquired. Seated right next to him, the Göttingen speaker sneered that the workmanship was inferior to what one found on the Continent. Not to be outdone, von Neumann slapped an ivory slide rule onto the table with algorithmic enthusiasm. Gödel, with characteristic gravity, raised a glass of port in a toast that seemed prepared for its own incompleteness. The scholar just back from Oxford preferred brandy and, being full of it, soon leapt onto the table to make a point that no one had invited. In the ensuing disorder, a fellow guest's black coffee went flying. That black coffee, in the left-to-right order of cups along the table, had been sitting somewhere between Hardy's espresso and Ramanujan's tea. By morning the hall was deserted. Under the table lay four instruments: the antique abacus, the brass astrolabe, the ivory slide rule, and the golden compass. The silver caliper was gone. Who possessed each instrument -- and who had been carrying the missing silver caliper?
Made with AI
1
4
12
2,078
Mayur Naik retweeted
MathDuels update: Gemini-3.5-Flash ranking #4, surpassing GPT-5.2 & Claude-Opus-4.7 Impressive performance for its speed & size.
2
7
524
Mayur Naik retweeted
nice math problem i came up with last december... Consider a cube of side length 2, aligned with the coordinate axes. Place three cylinders inside it, each of height 2 and radius R, each aligned with some coordinate axis. The cylinders may not intersect. What is the maximal R?
54
49
600
211,269
Incredibly proud of PhDs #8, #9, and #10 -- congratulations Drs. Jiani Huang (@jiani_huang_ai), Aaditya Naik (@aaditya_naik), and Adam Stein (@adamlsteinl)! I am lucky and grateful for the privilege of working with the three of you!
3
2
49
3,311
Mayur Naik retweeted
Very excited to share a new milestone in AI for Math: Aletheia, powered by Gemini Deep Think, was just used to autonomously solve a Kirby problem! “Kirby’s list” is a “compendium of the most important unsolved problems in topology, the study of deformable shapes” (Quanta magazine). 🧵
30
106
531
94,319
Mayur Naik retweeted
We are delighted to unveil our research blog Rabdology at rabdology.ai, where we chart the jagged math-frontier of AI reasoning. This is our first post in a weekly series. Read on, and if you enjoy it, please subscribe! (Link at bottom of blog's main page.) The Three-Cylinders Problem: When AI models choose Beauty over Truth rabdology.ai/three-cylinders We pose a problem that a good geometry student can solve in twenty minutes. We gave it to four of the world’s most advanced AI models and watched what happened. Three of them got it wrong — and the way they got it wrong tells you something different about the state of AI mathematical reasoning than the usual benchmarks.
1
7
15
1,731
Very timely, especially in light of revelation that 1/3rd of problems in FrontierMath are fatally flawed. As expert human validation of frontier math tasks approaches its inevitable limit, LLMs are stepping in to fill the void. But our work below shows that the discovery of mistakes in FrontierMath problems didn't necessarily have to wait for a frontier model like GPT-5.5. Much smaller and even open-source models can be as effective at verifying math proofs: their weights embody the necessary knowledge, as one might expect -- checking proofs ought to be easier than writing proofs. The crucial thing that makes this work is the use of "prompt ensembles", each of which modularly checks a different facet of a given proof, and some of which are even specific to the domain/sub-domain of math. Ideas like meta-prompting, agent skills, and autoresearch will undoubtedly evolve to make LLMs as judges of math proofs even more effective in future.
Do we need frontier models to verify math proofs? EpochAI just announced that they found several fatal flaws in their FrontierMath benchmark using GPT-5.5. But isn't verification supposed to be easier than generation, so why were they not spotted earlier? In our recent work, we asked a related question: do we really need frontier-scale compute to verify Olympiad-level math proofs? Turns out, even 20B open-source models can keep up with frontier LLMs on proof verification. Work done with my co-authors @aaditya_naik, @AI4Code, and @RajeevAlur Preprint: arxiv.org/abs/2604.02450
1
1
10
1,795
I feel sorry for either this person or their PhD advisor for them to say "The thing that still matters —> taste <— was never really taught". This is *most* of what my advisor taught me; the rest (semantics, coding, etc) was just hard work; he answered questions and gave feedback.
I was talking to a CS professor yesterday and he told me he's so relieved ⛱️ now that his PhD students have left for their summer internships.🌞. He finally gets to do interesting research himself again. It wasn't a complaint. Just something he'd been sitting with. The old PhD bargain was pretty clean: professors had ideas, students had time, execution was the bridge. You did the work, and somewhere along the way, taste rubbed off. That bridge is mostly gone now. A senior researcher with taste and a research agents ships faster than the same researcher plus a junior student. Explaining 🎙️ an idea to someone who doesn't yet have taste costs more than just building it. (the reason professor still needs students is because the research agent 🤖 is not yet capable enough) Nobody wants to say this out loud Professors keep mentoring because the system 🏫 asks them to, not because it's the highest-leverage thing they could be doing with their week. The thing that still matters —> taste <— was never really taught. It was supposed to leak out of the execution. Now there's no execution to leak from.
5
13
134
28,574
Mayur Naik retweeted
I've just finished a week where I have been using agentic workflows basically 24/7. Here are some impressions: 1. I am extremely tired. Supervision of so many projects is distracting, frustrating and exhausting. 2. I have never worked on so many fronts at once. It's not only a superficial impression. I've read a ton, did courses with students, coded, learned new mathematical theorems, explored new git repos. 3. The more you work with agents, the better you become at instructing them. Don't overthink the prompts. Doing things step by step and letting the agent guide you is very productive. 4. I failed in some of those projects. Basically I got the impression that I could do everything with the agent. Maybe, but time and tokens are constrained. 5. I don't use much ready made software anymore. I build or scaffold stuff that I need. My favorite modality is text and best interface is simple text API. Everything can be pipelined. 6. Mathematics can be treated as software. Whatever you need to do with it, it can be decomposed into small parts. The thinking bit is still mostly done in my brain if necessary. 7. I think of my past life as a warm-up for this new period. The offline reading, coding and internalizing was super important. Without it, I would be a pure vibe-coder. With training, you sail a ship. 8. I don't want to see my paychecks. It's expensive, because the better you get, the more you spend on more and more projects. You abandon them in a more advanced state, but there is a token price to pay. 9. Each time I see a challenge, I think whether I can find a solution with the agent. I don't skip tasks. Sometimes I feel overly optimistic. It's part of the learning to see where the agents are limited. 10. I think I master entirely new workflows. They are not properly described in the books. It's a brave new world. Each standard task can be turned into a task for the agent. It's very enabling. Happy coding everyone!
20
40
315
41,439
Mayur Naik retweeted
In MathDuels leaderboard, Gemini-3-Flash dominates at its size: the only models ahead of it are the latest, largest frontier releases. Gemma-4-31b-it has the highest author rating of any opensource model. Thanks @GoogleDeepMind for these smart little models 🥹
Static math benchmarks saturate. We built one that doesn't. Announcing MathDuels, the first self-play math benchmark. Every frontier LLM writes problems for the others, and is graded on the ones written for it. As models improve, so does the benchmark.
2
13
1,426
Mayur Naik retweeted
Excellent idea. A revival of mathematical duels but in the AI form. Which model is the best problem composer and which one solves problems with no trouble. And the judge is a human professional.
Static math benchmarks saturate. We built one that doesn't. Announcing MathDuels, the first self-play math benchmark. Every frontier LLM writes problems for the others, and is graded on the ones written for it. As models improve, so does the benchmark.
3
10
66
5,493