Matthew Johnson retweeted
Jax added support for opaque types with custom tangents. One cool use is proper handling of batched geometric objects. Here's a short demo of a differentiable rasterization, i.e. min L2 w/ 1500 triangles to a photo. Interesting to move beyond tensors. (docs.jax.dev/en/latest/hijax…)
3
11
178
21,130
Matthew Johnson retweeted
Just some personal thoughts now that the AI co-mathematician tech report is public... First, I'm so excited to see the co-mathematician team's hard work out for the world to preview. 💪+🦾=🔥 The team has built a system for mathematicians, with mathematicians. The fact it's now top of the FrontierMath leaderboard is a cherry on top, not the goal. Vibes and utility >> benchmarks. The system is currently being tested with a small number of professional mathematicians. It is not widely available, but I personally hope that, one day, we can get even more capable systems into the hands of all mathematicians. It's been a privilege working with this team at Google DeepMind since January. Props to @dhhzheng, @ADaviesAI, and @pushmeet for their leadership. Give them all a follow to not miss exciting upcoming work.
The future of Math is mathematicians and AI agents working together. Very pleased to introduce @GoogleDeepMind's AI co-mathematician: a multi-agent system designed to actively collaborate with human experts on open-ended research mathematics. Mathematicians testing the agent across areas as diverse as group theory, Hamiltonian systems, and algebraic combinatorics have reported impressive results. In autonomous mode evaluation on the rigorous FrontierMath Tier 4 problems, AI co-mathematician scored an unprecedented 48% — a new high score among all AI systems evaluated.
13
26
253
65,466
Matthew Johnson retweeted
Announcing Talkie: a new, open-weight historical LLM! We trained and finetuned a 13B model on a newly-curated dataset of only pre-1930 data. Try it below! with @AlecRad and @status_effects 🧵
197
453
3,637
1,467,032
Matthew Johnson retweeted
New work with @AlecRad and @DavidDuvenaud: Have you ever dreamed of talking to someone from the past? Introducing talkie, a 13B model trained only on pre-1931 text. Vintage models should help us to understand how LMs generalize (e.g., can we teach talkie to code?). Thread:
179
399
3,208
1,240,122
Matthew Johnson retweeted
GPT-5.2 derived a novel result in theoretical physics, showing that a type of particle interaction many physicists expected would not occur can in fact arise under specific conditions. There is great promise in the potential of AI to benefit people by accelerating science.
GPT-5.2 derived a new result in theoretical physics. We’re releasing the result in a preprint with researchers from @the_IAS, @VanderbiltU, @Cambridge_Uni, and @Harvard. It shows that a gluon interaction many physicists expected would not occur can arise under specific conditions. openai.com/index/new-result-…
197
216
2,343
409,448
Matthew Johnson retweeted
Very happy to be involved in the core training team. Really amazing to see JAX running on our ultra-large scale GPU clusters.
Since xAI was formed just 30 months ago, the small and talented team has made remarkable progress. The future has never looked more exciting!
7
6
126
5,708
Matthew Johnson retweeted
I have to mention this, this opus is reasonably good at low level jax, sharding and pallas. I would call it shard sherrif.
Introducing Claude Opus 4.6. Our smartest model got an upgrade. Opus 4.6 plans more carefully, sustains agentic tasks for longer, operates reliably in massive codebases, and catches its own mistakes. It’s also our first Opus-class model with 1M token context in beta.
4
3
65
5,891
Matthew Johnson retweeted
We frictionlessly trained on AMD GPUs and TPUs with a unified JAX framework. Our goodput for flagship runs went past 90%. @YashVanjani @mjcOhio @alokpathy @pcmonk painstakingly removed obstacles to maximize experimental velocity.
2
12
167
33,812
Matthew Johnson retweeted
Nano Banana Pro: "Generate a diagram of a two-layer neural network in the style of Stephen Biesty"
23
60
738
267,027
Matthew Johnson retweeted
Unbelievable: the famed Berkeley Math Circle is being forced to shut down due to a bureaucratic requirement where a guest lecturer giving an hour long lesson needs to be officially fingerprinted. How is fingerprinting even still a thing in the 21st century? Chancellor Lyons @richlyons: can you see the absurdity of the situation and figure out a solution? dailycal.org/news/campus/gen…
31
78
746
274,296
Matthew Johnson retweeted
SGLang now has a pure Jax backend, and it runs natively on TPU!
SGLang now runs natively on TPU with a new pure Jax backend! SGLang-Jax leverages SGLang's high-performance server architecture and uses Jax to compile the model's forward pass. By combining SGLang and Jax, it delivers fast, native TPU inference while maintaining support for advanced features like continuous batching, prefix caching, parallelism, speculative decoding, and highly optimized TPU kernels. Learn more in the blog below👇
2
5
154
21,806
Matthew Johnson retweeted
⛵Marin 32B Base (mantis) is done training! It is the best open-source base model (beating OLMo 2 32B Base) and it’s even close to the best comparably-sized open-weight base models, Gemma 3 27B PT and Qwen 2.5 32B Base. Ranking across 19 benchmarks:
20
87
594
127,818
Matthew Johnson retweeted
TPU-style collective matmuls on GPU!
Want to improve GPU compute/comms overlap? We just published a new short tutorial for you! A few small changes to the Pallas:MGPU matmul kernel is all it takes to turn it into an all-gather collective matmul that overlaps NVLINK comms with local compute: docs.jax.dev/en/latest/palla…
8
27
4,737
Matthew Johnson retweeted
Want to improve GPU compute/comms overlap? We just published a new short tutorial for you! A few small changes to the Pallas:MGPU matmul kernel is all it takes to turn it into an all-gather collective matmul that overlaps NVLINK comms with local compute: docs.jax.dev/en/latest/palla…
7
46
302
33,699
Matthew Johnson retweeted
Curious how to write SOTA performance Blackwell matmul kernels using MGPU? We just published a short step-by-step tutorial: docs.jax.dev/en/latest/palla… At each step, we show exactly what (small) changes are necessary to refine the kernel and the final kernel is just under 150 lines.
4
66
412
55,459
Matthew Johnson retweeted
Today we're putting out an update to the JAX TPU book, this time on GPUs. How do GPUs work, especially compared to TPUs? How are they networked? And how does this affect LLM training? 1/n
38
516
3,416
405,814
Matthew Johnson retweeted
So about a month ago, Percy posted a version of this plot of our Marin 32B pretraining run. We got a lot of feedback, both public and private, that the spikes were bad. (This is a thread about how we fixed the spikes. Bear with me. )
Marin 32B training crossed 1.5 trillion tokens today...
23
105
1,028
307,580
Matthew Johnson retweeted
Strong recommend for this book and the JAX/TPU docs, even if you are using Torch / GPUs. Clean notation and mental model for some challenging ideas. github.com/jax-ml/scaling-bo… github.com/jax-ml/scaling-bo… docs.jax.dev/en/latest/noteb…
9
159
1,148
77,792