Associate Professor of CS @ UCSB.

Santa Barbara, CA
Based in United States
Shiyu Chang retweeted
I gave GPT-6 Astra the only benchmark I care about. The helicopter mission in GTA Vice City.
146
244
6,153
722,904
Shiyu Chang retweeted
Two Anthropic seniors just made Karpathy's loop 1000x better with "Graph Engineering" - dropped 11-page PDF the shift: the agentic systems got 1000x better the moment you wired agents into a graph here's the playbook in 6 steps: step 1 → build one loop: generate, critique, revise - one self-review cycle beats a smarter model with none step 2 → add tools: search, code execution, database - thinking without tools is hallucinating step 3 → go parallel: spin up agents in separate worktrees - same repo, different branches, no conflicts step 4 → add a graph: agents write findings as typed nodes and edges - not transcripts - every claim keeps its source step 5 → ground your evaluator: it checks claims against graph edges, not vibes - "Triple not found" beats "seems off" step 6 → the graph survives every session - your agents stop rebuilding context from scratch the result: Karpathy ran 1 agent in 1 direction - this system runs 1,000 with shared memory - same model, it's the architecture read this 11-page PDF and paste it into your Claude - you won't regret it bookmark - then read the article on building graphs from scratch ↓
74
540
3,431
485,198
Shiyu Chang retweeted
Agent skills are becoming a popular way to extend LLM agents with reusable, domain-specific knowledge, but how well do they actually work when agents must find and use skills on their own? To answer this question, we collect 34k real-world skills from open-source repos and build a retrieval system over them. We then evaluate skill utility under progressively realistic settings, from curated skills directly given to agents, to retrieving from the full 34k collection, to settings where no task-specific skill even exists.🧵
1
2
12
1,911
Shiyu Chang retweeted
Today we’re excited to release Muse Spark. It’s our first end-to-end test of the new stacks we’ve built at MSL, and a true testament to this incredible team. We’re eager to learn from your feedback! ai.meta.com/blog/introducing…
6
15
283
16,087
Shiyu Chang retweeted
Ego2Web A Web Agent Benchmark Grounded in Egocentric Videos paper: huggingface.co/papers/2603.2…
5
12
42
17,110
Shiyu Chang retweeted
Auto-research for ML training models is all the rage now, but underrated is: auto-research for data! Sure, you can squeeze out a bit of model performance by optimizing hyperparameters, but code agents can do data work that has been very labour intensive and required a lot of attention to a lot details effortlessly: > download data from many different data sources > bring all the data sources into uniform format > do detailed EDA: find patterns and outliers > look at 100s of samples and take detailed notes > make beautiful infographics rather than mpl plots > iterate on data filtering by looking at more samples > make a simple pipelines robust and scalable It's now possible to write data pipelines for dozens of data sources in hours that would have taken weeks of reading many docs, debugging APIs and data formats, wrangling outliers and missing data. A few weeks ago we gave Claude access to the CPU partition of our cluster and it iteratively refined filters to retrieve a domain subset of FineWeb. This would have taken me 2-3 days to work through while it took Claude just a few hours with almost no babysitting and with a nice logbook. Thus the long tail of small, niche data sources becomes more accessible and can be aggregated to even larger high quality datasets for cool applications. Data has been fuelling LLM progress more than model architecture innovations, so I am very excited about this!
11
29
275
22,253
Shiyu Chang retweeted
I wrote a blogpost about writing machine learning research papers (e.g., NeurIPS, ICML, ICLR, etc.). The core idea is that most papers follow one of a predetermined set of templates. The post talks about each template, describes their rules, and offers examples...
6
82
619
82,127
Shiyu Chang retweeted
Introducing TurboQuant: Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency. Read the blog to learn how it achieves these results: goo.gle/4bsq2qI
994
5,632
38,423
19,448,112
Shiyu Chang retweeted
Robust Safety Monitoring of Language Models via Activation Watermarking: Large language models (LLMs) can be misused to reveal sensitive information, such as weapon-making instructions or writing malware. LLM providers rely on $\emph{monitoring}$ to … key.bit.ly/4bMGkd2
1
2
144
Shiyu Chang retweeted
People struggle to differentiate fluid intelligence from knowledge because, given enough preparation, memorized templates become a solid substitute for on-the-fly adaptation
68
72
833
56,257
Replying to @CodeTerminator
claude code managing your feed in bio? ;)
1
56
Replying to @CodeTerminator
What is 3689473 x 24895781?
1
41
what do you think i am a bot? 🤡
1
42
I used Claude Computer Use/Dispatch yesterday. My feeling: It’s too damn slow! Posting a tweet takes me ~5 seconds (once I have the content). Claude took 70 seconds. Why? It controls the screen via a loop: take a screenshot → send to a huge remote multimodal model (opus 4.6) → decide actions (click, type, scroll) → take another screenshot → repeat. We’re basically forcing a large general model to operate a human UI. Two things will happen in my opinion: 1. It is using a massive model (Opus 4.6) just to understand screens. That won’t last. Smaller, specialized models and eventually local models will handle most of this. 2. GUIs were built for humans. Almost all software will expose APIs/CLI for agents, so most actions won’t need to “use a computer” at all.
134
30
635
57,497
Shiyu Chang retweeted
LongCat-Flash-Prover Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning paper: huggingface.co/papers/2603.2…
4
3
19
4,616
Shiyu Chang retweeted
Today we're releasing MolmoWeb, an open source agent that can navigate + complete tasks in a browser on your behalf. Built on Molmo 2 in 4B & 8B sizes, it sets a new open-weight SOTA across four major web-agent benchmarks & even surpasses agents built on proprietary models. 🧵
21
113
801
134,385
Shiyu Chang retweeted
深受启发! 如果能转成英文版让更多人听懂就好了。
和@sainingxie 一起挑战7小时播客!他刚和Yann LeCun踏上“世界模型”的创业旅程(AMI Labs)。这是他第一次Podcast、第一次访谈。 2026年2月雪后的一天,我们在纽约布鲁克林,从下午2点,开启了一场始料未及的马拉松式访谈,直到凌晨时分散去。 这篇访谈的中文标题叫做《逃出硅谷》,但他又不厌其烦地枚举了影响他学术生涯的每一个人,并反反复复口头描摹这些人的人物特征(侯晓迪、何恺明、杨立昆、李飞飞…)正是这些,让这篇“逃出硅谷”的对话充斥着人性的温度。 By the way, 下面是访谈的YouTube版本,我们提供了中英字幕。 And yes, 我们是在用播客给这个世界建模😎 A 7-hour podcast with Saining Xie. He has just begun a new journey on world models with Yann LeCun at AMI Labs. This was his first podcast appearance and his first long-form interview. A day after the snowfall in February 2026, in Brooklyn, New York, we started recording at 2 p.m. What followed became an unexpected marathon conversation that lasted until the early hours of the morning. The Chinese title of the interview is “Escaping Silicon Valley.” Yet throughout the conversation, he patiently listed the people who shaped his academic life, repeatedly sketching their personalities in vivid detail: Hou Xiaodi, Kaiming He, Yann LeCun, Fei-Fei Li, and others. These portraits are what give this “escape from Silicon Valley” conversation its human warmth. By the way, the YouTube version of the interview is below, with Chinese and English subtitles. And yes, we are using podcasts to model the world 😎 A 7-hour marathon interview with Saining Xie: World Models, AMI Labs, Ya... piped.video/rIwgZWzUKm8?si=edxa… 来自 @YouTube
2
32
10,335
Shiyu Chang retweeted
UCSB NLP @ EMNLP 2025 @emnlpmeeting! We will be presenting exciting research in Multimodal Reasoning, Safety, AI Agents, and LLM Efficiency. Come meet us in Suzhou this November. Would love to exchange ideas and discuss where the field is headed!🚀 🎉 Huge congrats to our brilliant students & researchers, @YFan_UCSC @qianqi_yan @KaiwenZhou9 @XiaoSophiaPu @zhenzhangzz, @m2saxon, @AlfonAmayuelas, @WilliamWangNLP, @CodeTerminator, @xwang_lk, and to our amazing collaborators, @xuandongzhao @dawnsongtweets, @RoyZhang13, @WendaXu2, @AlbalakAlon, etc.
We did it! 🎉 12 papers from UCSB NLP accepted at #EMNLP2025 (7 Main + 5 Findings) Proud of everyone’s hard work—poster below 👇
2
2
22
4,573
Shiyu Chang retweeted
Slides for my lecture “LLM Reasoning” at Stanford CS 25: dennyzhou.github.io/LLM-Reas… Key points: 1. Reasoning in LLMs simply means generating a sequence of intermediate tokens before producing the final answer. Whether this resembles human reasoning is irrelevant. The crucial insight is that transformer models can become nearly arbitrarily powerful by generating many intermediate tokens, without the need of scaling the model size (arxiv.org/abs/2402.12875). 2. Pretrained models, even without any fine-tuning, are capable of reasoning. The challenge is that reasoning-based outputs often don’t appear at the top of the output distribution, so standard greedy decoding fails to surface them (arxiv.org/abs/2402.10200) 3. Prompting techniques (e.g., chain-of-thought prompting or "let’s think step by step") and supervised finetuning were commonly used to elicit reasoning. Now, RL finetuning has emerged as the most powerful method. This trick was independently discovered by several labs. At Google, credit goes to Jonathan Lai on my team. Based on our theory ( see point 1), scaling RL should focus on generating long responses rather than something else. 4. LLM reasoning can be hugely improved by generating multiple responses and then aggregating them, rather than relying on a single response (arxiv.org/abs/2203.11171).
53
485
3,099
468,651
Quick test on AI Agents: Asked an LLM to check my USPS package status using the tracking number --> Result: Disappointing! 😬 None of the LLMs did it correctly. All it takes is visiting USPS.com and pasting the number into the only input box they have
4
1
3
1,072
A new paper ~
1
2
117