reducing perplexity & ai for science | past: reasoning o1/o3/... @openai & probabilistic programs, proteins, science & reasoning @ google brain đź§ 

David Dohan retweeted
Not included here, but worth saying: modeling ourselves as in an “AI race” really ceases to make any sense immediately before RSI. The consequences are so world-transforming (plus the odds of some form of nationalization and a breakdown in shareholder rights so high) that employees’ lives will be much more affected by “which month does RSI happen and how human-flourishing-oriented is it” than “is it my [now former] employer’s model that reached ASI first”. Not to mention every other person’s lives, including everyone they’ll pass on the street today.
5
10
77
3,969
David Dohan retweeted
yes, nonsofic groups exist: this statement is one of many new beautiful results proved by Astra, our next major model. We're releasing 10 such Astra proofs, complete with lean certificates and CoT walkthroughs for each of them. The results are wide-ranging, from von Neumann algebras (disproof of Connes' Rigidity Conjecture) to better bounds for high dimensional sphere packing, for circuit complexity, for monochromatic triangles in multicolored graphs, and more. More thoughts here: openai.com/index/ten-advance…
277
944
6,676
4,280,516
David Dohan retweeted
Smalltalk and HyperCard: often mentioned together, but surprisingly different design philosophies IMO! The Smalltalk school: everything is the same thing. Minimal, elegant. If you want to mod the graphics system be my guest. Powerful but a bit scary and hard to get started in my own experience. The HyperCard school: it’s just a friendly drawing app, don’t worry :) oh and you can add links. Oh and you can add buttons. Oh and you can customize the button scripts. Oh whoa you made a game! Both have benefits, neither is more “right”. Arguably some of their virtues may even be orthogonal. But I think the smartest programmers tend to underestimate the value of the HyperCard approach.
11
7
124
11,734
David Dohan retweeted
I was always slow on Forward v.s. Reverse KL. Then I realized a mnemonic trick. FOrward -> FOMO Fear of missing out on some probability mass: min_q KL(p || q) encourages mode coverage of p. Reverse -> Retreat Retreat to a high-density mode: min_q KL(q || p) is mode seeking.
8
19
252
24,006
David Dohan retweeted
AI researcher: *through sobs* you can't just say everything is conscious.... Please.... Me: *points at rock* conscious
29
100
1,453
David Dohan retweeted
A workflow I'm enjoying: "Walk-driven development" > go on a nice walk outside 🚶 > record a long audio note: ideas, goals, things to build 🎙️ > agent auto-creates docs/tasks, and kicks off cloud coding agents for me 🤖
48
47
588
77,098
Now imagine it only takes 2 years
Imagine what it will be like if 5 years from now models have improved on Fable as much as Fable has improved on GPT3.
7
4
167
19,097
Big if true, and afaict it's true congrats to the Etched team on getting inference running on their new chips!
Replying to @Etched
Introducing Low-Voltage Inference (LVI) for high throughput workloads. Today, AI chips can't scale FLOPs without thermal throttling. As FLOPs utilization increases, AI chips draw more power and downregulate clock speed. This often results in sustained inference throughput under half of peak FLOPs. Chips in other industries solve the power problem by running at lower voltages. Bitcoin miners run at under 3x the voltage of AI chips! We’ve designed a new architecture to run our chip’s math blocks at under half the voltage of most AI chips. This enables multiple times the FLOPs density of AI chips today. We can run trillion parameter sparse MoEs at 80%+ peak FLOPs without thermal throttling. Running LVI requires co-designing the entire cluster from the transistor to the token: new splittable math arrays, circuit techniques, novel tiling and scheduling algorithms, power delivery networks, VRM architectures, advanced packaging, cold plate designs, and more.
40
6,973
David Dohan retweeted
We achieved gold medal-level performance 🥇on the 2025 International Mathematical Olympiad with a general-purpose reasoning LLM! Our model solved world-class math problems—at the level of top human contestants. A major milestone for AI and mathematics.
1/N I’m excited to share that our latest @OpenAI experimental reasoning LLM has achieved a longstanding grand challenge in AI: gold medal-level performance on the world’s most prestigious math competition—the International Math Olympiad (IMO).
214
421
3,958
674,936
David Dohan retweeted
I feel this may be helpful to some of you today:
13
60
704
89,262
Fun to watch prediction markets update on the news
1
15
2,101
OpenAI achieved gold medal on 2025 International Math Olympiad (solving 5 of 6 problems)! Thinks for hours and writes proofs in natural language. We've come a long way from LLMs solving 50% of MATH dataset in 2022 Congrats @alexwei_ on spearheading a major milestone!
1/N I’m excited to share that our latest @OpenAI experimental reasoning LLM has achieved a longstanding grand challenge in AI: gold medal-level performance on the world’s most prestigious math competition—the International Math Olympiad (IMO).
1
126
7,764
How to code a side project in 2025: 1. May 31 - Write project spec 2. Procrastinate 6 months 3. Dec 31 - ask favorite AI to implement it
2
3
66
4,542
David Dohan retweeted
Scaling pretraining and scaling thinking are two different dimensions of improvement. They are complementary, not in competition.
48
80
1,020
129,504
David Dohan retweeted
This is on the scale of the Apollo Program and Manhattan Project when measured as a fraction of GDP. This kind of investment only happens when the science is carefully vetted and people believe it will succeed and be completely transformative. I agree it’s the right time.
Announcing The Stargate Project The Stargate Project is a new company which intends to invest $500 billion over the next four years building new AI infrastructure for OpenAI in the United States. We will begin deploying $100 billion immediately. This infrastructure will secure American leadership in AI, create hundreds of thousands of American jobs, and generate massive economic benefit for the entire world. This project will not only support the re-industrialization of the United States but also provide a strategic capability to protect the national security of America and its allies. The initial equity funders in Stargate are SoftBank, OpenAI, Oracle, and MGX. SoftBank and OpenAI are the lead partners for Stargate, with SoftBank having financial responsibility and OpenAI having operational responsibility. Masayoshi Son will be the chairman. Arm, Microsoft, NVIDIA, Oracle, and OpenAI are the key initial technology partners. The buildout is currently underway, starting in Texas, and we are evaluating potential sites across the country for more campuses as we finalize definitive agreements. As part of Stargate, Oracle, NVIDIA, and OpenAI will closely collaborate to build and operate this computing system. This builds on a deep collaboration between OpenAI and NVIDIA going back to 2016 and a newer partnership between OpenAI and Oracle. This also builds on the existing OpenAI partnership with Microsoft. OpenAI will continue to increase its consumption of Azure as OpenAI continues its work with Microsoft with this additional compute to train leading models and deliver great products and services. All of us look forward to continuing to build and develop AI—and in particular AGI—for the benefit of all of humanity. We believe that this new step is critical on the path, and will enable creative people to figure out how to use AI to elevate humanity.
243
658
7,308
918,253
David Dohan retweeted
🚨SCANDAL 🚨 OpenAI trained on the train set for the Millenium Puzzles
81
27
1,627
142,777
David Dohan retweeted
these new captchas are getting way too difficult
28
93
2,010
77,408
David Dohan retweeted
o3 has literally made 0% progress on the Millennium eval it’s ai winter now
54
26
1,761
191,802
David Dohan retweeted
I have yet to find a well-defined task that cannot be optimized by these models. Eval improvement like ARC AGI showcase this dynamic
So we went from 0 to 87% in 5 years in ARC AGI score. There is no wall it seems. GPT-2 (2019): 0% GPT-3 (2020): 0% GPT-4 (2023): 2% GPT-4o (2024): 5% o1-preview (2024): 21% o1 high (2024): 32% o1 Pro (2024): ~50% o3 tuned low (2024): 76% o3 tuned high (2024): 87%
6
8
110
32,318