happy to announce that we've gotten rid of tokenizers! especially excited with what we've replaced them with: end-to-end trainable modules that not only learn to group characters into (sub)words, but can iterate to group words into phrases and further higher-order concepts see @sukjun_hwang's thread for more details 👇
Tokenization has been the final barrier to truly end-to-end language models. We developed the H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data
13
50
759
82,013
absolutely incredible
Green hydrogen currently runs on iridium — one of Earth's rarest metals, a few tonnes mined per year. Our AI-directed lab screened 2,942 catalysts in 3 months and surfaced a palladium-based lead the field had written off, now past 1,000 hours of testing in acid. lila.ai/news/how-an-ai-run-l…
3
193
14,582
brandon wang retweeted
Without a mechanism, I’m struggling to understand the importance of this. Host-pathogen interactions have been a rich source of biological tools for decades; it’s where CRISPR and restriction enzymes and Agrobacterium came from. But the way we get these tools to be great, and useful, is by understanding their mechanism such that we can begin engineering and improving them. I’m sure AI tools will be helpful for molecular characterization and engineering; so why not wait to announce until you actually do that?
We’ve set up a molecular biology lab at Anthropic and we’re announcing our first discovery! Claude discovered a new CRISPR-like enzyme. 950 agents spent 21 hours searching through a database of DNA sequences until one of the agents found something striking: “[The DNA next to the RT] is spectacular: I can see by eye a tandem repeat array … that's a CRISPR-like … repeat array?!”. After analysis and testing in our lab, we found that the sequence is a previously uncharacterized enzyme system. We don’t know what it does yet, but it has features reminiscent of CRISPR. Our lab looks like a typical molecular biology lab. Our research only involves the lower-levels of biosafety risk level; we don’t handle pathogens that can infect humans, and all the lab work is performed by human scientists. We’re sharing these early findings with the community to show how Claude can be used to accelerate fundamental research in biology.
15
33
324
18,575
i dont get why people are reading so much into muse downloads/comparing it to chat etc. threads did similar numbers meta gets inf downloads on anything they make since they can just put ads that show up whenever you open ig
5
260
brandon wang retweeted
I think there is an urgent need for every community to declare and commit to a central core value: we (humans) should help humanity flourish. Thanks to Terry Tao for highlighting many different views on the future of #Math+#AI on his blog by cross-posting guest posts: terrytao.wordpress.com/2026/… More interactive version: poshenloh.com/posts/20260919… The post reasons through why that human commitment then reconciles several conflicting positions that different groups had been taking over recent months: (1) it's important to preserve a strong community of human experts in every line of work (not just math), (2) AI can create significant new value for humanity, (3) people should change with technology, and (4) AI is going too fast, and must to slow down. Along the way, I predict that AI will slow down because we ironically have just created ... too many jobs to fill. That's because of the specter of AI-accelerated hacking. I hope this can lead more (human) communities to publicly declare that they are organizing their AI-driven strategic changes to advance human flourishing. Personal note: Ever since I had the opportunity to take a class from Terry Tao back at UCLA in 2008 (as a not-particularly noteworthy student), I was struck by his humility amidst astonishing brilliance (not just in math, but remarkably broadly). So, I knew that when I saw his name among the 25-Fields-Medalist declaration about the misalignment between AI and Math, he was trying to act in good faith. I am thrilled to collaborate in this way.
18
64
349
32,082
a lot of the surprise of ai x math advances comes from people not internalizing test time compute scaling properly. inference compute per problem has gone up sharply this year (consumer gpt in jan, $2k for unit distance, $15m for ns) partly bc ttc scaling has improved a lot
1
8
550
wow claude for finance is getting pretty sophisticated
Straits Taylor Rule: i = r* + π* + 1.5(π−π*) + 0.5(y−y*) + α(SOH−SOH*) + β(BEM−BEM*), α,β > 0 Let’s see if a hike could open SOH or produce a single barrel :) You can’t 25bp a chokepoint and r* isn’t neutral. It’s SOH risk premium, and We set it. Stay unanchored !
1
452
this is super cool
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/
3
777
brandon wang retweeted
Dario deserves tremendous credit for pushing the industry forward on AI safety. But his approach to China is fundamentally flawed. Cooperation on AI safety is not a negotiation where the US uses leverage to bring China to the table. That approach will backfire, amplifying China’s distrust of the US on AI and making it harder to find areas of productive exchange. There’s a limit to how much the US can change China’s behavior through power alone. In the end, we will need China to want to cooperate and address AI risk in a real way.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
40
51
255
91,158
confusing to see people cite dsv4.1 as evidence that rl algorithms don't matter, even when they say "at the current stage." there's no doubt ds will revisit grpo in due time.
2
10
628
one of ds's biggest strengths is that each model iteration focuses on the most critical bottleneck of their last model, instead of making some dogmatic commitment to an arbitrary research bet or pushing on whatever the direction du jour is
1
6
321
in fairness, openai is usually quite good about putting out explainers, well-written manuscripts, getting comments from experts, etc. and it sounds like they'll still do that just unfortunate what happened with navier stokes (and also still confusing why it was so rushed out)
very disappointing to see this behavior, and also the rush to announce when tristan and co clearly delayed their release to properly flesh out the details in a more contributory manner cims.nyu.edu/~tristanb/state…
2
11
822
istg deepseek architecture always comes up with the funniest stuff
5
325
The news today of progress on resolving the Navier–Stokes problem, one of mathematics’ great longstanding challenges concerning the equations that govern the flow of fluids, represents a milestone advance in human knowledge. This story began with Navier, Stokes, Leray, and Ladyzhenskaya and has culminated in the recent breakthroughs of Córdoba and Martínez-Zoroa, then — assisted by new technologies — Alpöge and Buckmaster, with the final steps taken by OpenAI mathematicians. The purpose of mathematics is human understanding, and this achievement, and the process that led to it, will bear fruit for a long time to come. Ravi Vakil, President of the AMS, and John Meier, CEO of the AMS Read more. Link in comments.
61
708
2,675
475,229
very disappointing to see this behavior, and also the rush to announce when tristan and co clearly delayed their release to properly flesh out the details in a more contributory manner cims.nyu.edu/~tristanb/state…
3
2
63
3,159
beautiful
Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: anthropic.com/research/forma… And see the complete proof on GitHub: github.com/anthropics/fermat…
1
2
24
2,458
sorry what the hell tbh i'm surprised this has not solved more open problems jesus
Replying to @RyanGreenblatt
Ryan have you seen this? Astra (NONE) scores higher than 5.6 Sol Pro (Max) on FrontierMath T4
2
17
3,521
brandon wang retweeted
you better keep your promises tomorrow mr altman
GPT-6 will be renamed GPT-6-7, you're welcome
1
7
404
14,470
brandon wang retweeted
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
276
510
6,443
1,644,641
brandon wang retweeted
I've seen idea floating around that is something like "trading firm spending on frontier tokens is enough to justify massive capex from model providers"... What?! Trading is such a tiny fraction of the economy, more notable for PNL/person than any absolute number.
3
4
219
17,351
this was by far the best place I've been to in SF
10
1
74
64,599