ml @reductoai | formerly @Berkeley_AI | other things i like: music, econ, cooking, the #calgorithm, wikipedia, funny reels, not necessarily in that order

San Francisco, CA
Wrote a deep dive on implementing a language model from scratch in JAX and scaling it with distributed training! If you’re coming from PyTorch and want to see how the same ideas look in JAX, or just want a hands-on intro to distributed training, check out this blog post: chuyishang.com/blog/2026/jax… Comes with code + an assignment and test cases so you can follow along!
8
64
601
34,988
Very interesting idea and great paper!
There’s lots of buzz around agentic harnesses for robots, particularly for long horizon tasks that require complex reasoning and memory. But what will it really take to turn a reasoning agent into a reactive, reliable and low-latency robot policy? In our new paper, workspace models, we design a new memory architecture that acts as a latent harness for stronger reasoning models. Here’s why we think this might be the way forward: (1/9)
6
299
the tosh era begins 💪💪💪
POWER. OF. UNIT. We’ve got maximum capacity in Memorial Stadium tonight. See you at 7:30 🫵 #GoBears
1
5
195
pain 😢
2
39
chuyi shang retweeted
Excited to share our work on a model for mapping molecules to olfactory receptors to perception, out today! Check out the link in thread ⬇️
Out today in Cell Systems: the final version of our work on biologically-inspired models for predicting how molecules smell! With @SeyoneC and @jdthamores
6
16
82
8,358
had a lot of fun working on this model!! try it out -- we think you'll like it as well!
Today we’re announcing r-1, our new document parsing model. It’s more accurate than our most powerful agentic OCR models, faster, and up to 6x cheaper. At @reductoai, we spent two years building specialized models for complex visual layouts, tables spanning multiple pages, and key formatting like strikethroughs. We then used everything we learned to build r-1, the first in a new generation of models designed to handle the hardest documents without multiplying cost. This early preview delivers a 20% lower error rate than our most accurate legacy agentic models and will keep improving with new checkpoints over the next few weeks. Accuracy is only the beginning. r-1 preview is available at 1¢ per page all-in, with additional volume discounts as you scale. In the near term we’re also going to release r-1 mini, and an auto mode that intelligently selects the right approach for each page. If you’re currently using another parser, we’re offering up to $5,000 in credits to evaluate and migrate. You can claim the migration offer and learn more about r-1 using the links in the comments. Happy parsing!
1
26
939
chuyi shang retweeted
visit your alma mater. love this place. go bears!
new uc berkeley data science building
23
12
647
119,467
chuyi shang retweeted
Deciding when to jump in and help someone—and when to hold back and let them work through it—is something humans navigate constantly. How do AI assistants handle this tradeoff? We introduce Int-Bench, a framework for evaluating interventions during problem-solving tasks.
4
6
31
6,747
chuyi shang retweeted
Everyone is talking about memory lately: Micron, SanDisk, etc. Here, we zoom out from FlashAttention/device memory to the next bottleneck: data-center communication. That is where photonics matters. winterrykim.github.io/blog/2… w/@punhojark If interested, come see our work at ICML.
1
4
20
1,510
chuyi shang retweeted
Visual deep dive on FlashAttention by hand ✍️ (drawn with Excalidraw) winterrykim.github.io/blog/2…
4
37
297
12,150
Presenting this work today at @CVPR from 11:45 AM – 1:45 PM, Poster 451! Swing by if interested, would love to chat! #CVPR2026
What if the best visual reasoning steps are ones humans can’t specify? 🤔 Existing VLM reasoning is often constrained by language, pixels, and human-designed intermediates. We introduce Latent Implicit Visual Reasoning, where we show that VLMs can discover the best visual reasoning steps by themselves — no bboxes, no intermediate images, no extra supervision. Presenting this week at CVPR! (1/n)🧵
1
4
12
2,385
What if the best visual reasoning steps are ones humans can’t specify? 🤔 Existing VLM reasoning is often constrained by language, pixels, and human-designed intermediates. We introduce Latent Implicit Visual Reasoning, where we show that VLMs can discover the best visual reasoning steps by themselves — no bboxes, no intermediate images, no extra supervision. Presenting this week at CVPR! (1/n)🧵
2
9
34
8,893
Visual reasoning does not need to be hand-designed! By changing the training dynamics, we can let VLMs discover their own visual abstractions for solving multimodal tasks. 📝Paper: arxiv.org/pdf/2512.21218 🌐Website: chuyishang.com/livr/ 🧑‍💻Code/Data: Releasing later this week! Work with amazing collaborators @kelvinli01 @roeiherzig @leokarlin @RogerioFeris and @trevordarrell
1
7
8,795