Creator of TanML | Agentic AI Systems | Senior Quantitative Modeler (AI/ML)

Looking forward to speaking at @aiDotEngineer New York, Oct 12–14. Come say hi — ai.engineer/nyc/2026
1
2
670
Tanmay Sah, PhD retweeted
can confirm. ran @latentspacepod AINews side by side with 6 Sol and the difference was night and day: latent.space/p/ainews-claude… 5.5 Opus is the new default model for AINews going forward. so much more concise and tasteful reporting, with much less slopese than even 5 Opus.
also important news we fixed the writing
154
13
217
27,093
Coining “EvoUndo Engineering”: building agents that can safely undo their changes.
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses Tanmay Sah, Dolly Sah, Harshul Jain, Tanya Sah arxiv.org/abs/2608.28363 [𝚌𝚜.𝙰𝙸]
46
Tanmay Sah, PhD retweeted
part of this is in how ~everyone with a platform here speaks about LLMs for the past few months i rely on LLMs as excitedly as anyone, but... are we using the same models, guys?! the frontier is still dumb as a brick 20%+ of the time and needs more hand-holding than a freshman it's just that for some VERY specific types of work, like all the beautiful discovery announcements in the recent past, this hand-holding was painstakingly and expensively done by the labs for us (fantastic btw) but... 99% of the time, i am not searching under the lamppost of the AI lab's over-optimized workflows, so that high cost of alignment to the task is mine to handle and i absolutely do it because it's extremely valuable, and if it needs to be said, i think the progress has been far beyond what i'd have guessed and it's only gonna get far better from here. but don't let the jagged frontier of capabilities delude you into thinking that things are plainly "superhuman" at any *coherently broad* set of capabilities, because A LOT remains to be done (and i think it will be)
The vibe and language here have been drifting wholesale away from sensible, nuanced takes on the technical topics that animate me. Been oscillating on whether I should essentially just check out other than amplifying releases, as I've done for the bulk of this year to date. Versus forcing myself to articulate some of this in the long-form style of my 2023-2025 rants. So far, our collective goldfish memory here doesn't help make it feel worthwhile.
30
33
354
30,872
Tanmay Sah, PhD retweeted
Andrej Karpathy just explained the 5 shifts turning LLMs into agentic systems. 00:00 - Memory turns chat into personal AI 06:41 - Multimodal AI reads the world 16:58 - Thinking models solve harder tasks 24:51 - Search makes LLMs live 30:58 - Tools turn LLMs into workers Most people are still treating LLMs like chatbots. Karpathy is showing the full stack: Memory → Vision → Reasoning → Search → Tools Prompting is the old workflow. Agentic systems are the new one. This 40-minute talk is worth more than most paid AI agent courses. Bookmark and watch it before everyone catches up. Then read how to turn LLMs into self-improving agent loops below
36
466
2,792
355,871
The ACM CAIS 2026 session recordings are now live. Great to see the talks from San Jose available online, including our presentation on The Verifier Tax. Check out the full playlist below.
CAIS 2026 session recordings are live!: LOADS of talks from San Jose, now up on ACM's YouTube channel. Watch the full playlist: piped.video/playlist?list=PL…
2
50
If you missed the ACM CAIS 2026 conference and want to know which paper won the Best Paper Award, check out the YouTube video piped.video/watch?v=-vnH1WkQ… of their presentation. Great work! Shuren Xia Jorge Ortiz and the entire team!
46
Tanmay Sah, PhD retweeted
Anthropic pays $750,000+ a year for engineers who can build LLM architectures from scratch. Stanford taught the entire thing in 1 hour lecture & released it for free. Bookmark & watch this today before someone takes it down and read this article "How claude uses context"
6
25
108
25,028
Sometimes, a GIF speaks louder than words: consistency compounds. What do you think?
28
Tanmay Sah, PhD retweeted
Tomorrow I’ll show how to unify your data, services, and systems into a single virtual filesystem for AI agents using Mirage github.com/strukto-ai/mirage Instead of building custom integrations for every tool, expose everything through a familiar filesystem interface. Join if you’re curious about the future of agent system. #AIAgents #AgentSystem #VirtualFilesystem #Mirage
Less preaching, more practice. This Thursday three founders show what their in-production agents can do. An overnight software factory, graph memory, a whole company in one terminal. Live Q&A after every demo.
3
1
10
1,018
Tanmay Sah, PhD retweeted
Thanks for discussing meta-harness @lilianweng ! A very clearly written summary of the state of harness engineering + RSI. I often get the question"what can a harness actually improve?" which i only had partial answers to; this post lays out the whole picture in one place
new post on harness engineering for AI self-improvement: lilianweng.github.io/posts/2… It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple. Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
3
11
115
10,345
The Verifier Tax: Horizon Dependent Safety--Success Tradeoffs in Tool Using LLM Agents | Proceedings of the ACM Conference on AI and Agentic Systems dl.acm.org/doi/full/10.1145/…
46
Tanmay Sah, PhD retweeted
TIL verifier tax: an AI agent can complete a task and still fail. It might get the right outcome while breaking a policy, skipping an approval step, or exposing sensitive data. The task is done, but the process wasn't safe. Counts as success or failure? dl.acm.org/doi/full/10.1145/…
1
1
161
Tanmay Sah, PhD retweeted
Ty @heathercmiller for organizing the amazing @CAISconf event and @trq212 for the amazing talk! Had hella fun this week 🎉🎉🎉
3
27
1,729
Tanmay Sah, PhD retweeted
Really happy to share that our paper -“Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain” has won Best Paper (Outstanding Problem Paper) at @CAISconf 🏆 TL;DR - Poison ~2% of an AI agent’s fine-tuning traces and you can plant a trigger-activated backdoor that leaks confidential data >80% of the time. Guardrails completely miss it. (and thanks GPT for editing me into the team photo so cleanly nobody can tell I wasn’t actually there 😅)
Made with AI
2
6
13
1,731
Tanmay Sah, PhD retweeted
@CAISconf was sold out but small, focused on a particular well-scoped area (compound agentic systems). Some workshop papers were submitted as late as just a couple weeks before (fresh content), and good thing, too! Agentic systems already look quite different today than even 3 months ago. Perhaps the systems focus naturally attracts researchers who are inclined to build prototypes and libraries (not just publish), but I was still pleasantly surprised to see how many members of the Laude community had independently identified this conference as one they wanted to submit to and attend. On top of that, throwing a lounge after hours with food, drinks, demos, salon-style conversation, (and a mini podcast studio, because why not?) made it easy to get an even higher concentration of interesting people with interesting half-baked ideas in a particular niche that happens to be on fire right now. Loved the experience. Well-done, CAIS organizers (@heathercmiller, @lateinteraction, @matei_zaharia, @deeptir18, et al). Same time, same place next year?
1
3
15
4,653
One problem, two solutions, one through-line: agentic systems need real guarantees, not heuristics. All three, plus 60 more papers: caisconf.org/program/2026/pa…
1
3
301
Excited to share that our paper, The Verifier Tax: Safety–Success Tradeoffs in LLM Agents, will appear at @CAISconf 2026 in San Jose, where we'll also present the work. alphaxiv.org/abs/2603.19328 Thanks to the reviewers, @lateinteraction, @heathercmiller, and @aviaviavi__ 🙏
114
Most ML models do not fail in training. They fail in validation. As ML systems scale, validation is becoming the real bottleneck. Speaking this Friday, April 24, at the NC State MFM Workshop on Quantitative Finance on how TanML approaches this. #MachineLearning #MLOps #AI
119