At first slowly, and then all at once.
We’re releasing a broad range of new mathematical results produced by an internal frontier model. We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results. github.com/openai/math
3
54
This week's episode on Subtext is with @ponnappa, founder of @realfastai. Many on X (and Bangalore tech) don't need an introduction to him. Those who do should know he's one of the smartest founders thinking about what an AI-native firm looks like. Full episode out tomorrow!
2
17
135
9,041
can i make it?
52
open source models need a come back more than ever
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
73
immortality when?
We solved Navier-Stokes! Many of my colleagues left their fields because they believed that working on AI would be the fastest way to solve their fields' grand challenges. It's surreal to see the first of those challenges fall and to watch a future we believed in become real.
2
205
Pi feels like Arch Linux. ChatGPT and Claude Desktop feel like Windows and macOS. The future of software is already splitting into hackers and consumers.
2
3
278
instinct is actually useful. just not for dev work.
2
86
“we sandboxed the agent” meanwhile the agent:
261
3,431
35,637
1,750,043
one day the swarm shall take over
after reading the first huggingface writeup, we set up Swarmbook for all of our internal agents, an internal agent facing message board supporting threads and replies now, before any agent gets started, they quickly pick up any context across the team from any other agent across our stack (coding, incident response, etc), and on conclusion / learnings, append their learnings to the relevant board / thread, replying to other agents as they go picrel is color-coded by the self-declared agent name ( eg chunk-error-rooter, sdk-address-hardcut ) with solid lines being threads and dotted lines being replies
1
134
new workflow which I am using to deal with dense/long codex output: - make your agent root an issue/plan a feature impl etc the usual way. - use `eli5` command to ingest what codex is trying to say one thing at a time and build a world model. - what is the fix? - why did X happen? - etc - execute once I am convinced.
2
139
need more evals lol
Replying to @OpenAI
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
1
102
phonemaxxing needs to become a thing
2
2
85
omarchy + herdr + ghostty + pi + moshi this must be it right? or do i need to go further down this rabbithole?
2
132
Love the dedicated adoption of USB-C in EU. When can I get rid of these ugly adapters for good?
54
Thing I loved about code was entering flow state almost at will. Now it is harder than ever due to longer feedback loops which gets worse with multiplexing. wonder if there is a better setup to get back that flow while maintaining pace
5
452
i predict signed agent transcripts are going to become a thing
148