I work on building our AI coding experiences @Microsoft. Accidentally co-created GitHub spec-kit a few months ago (github.com/github/spec-kit).

Redmond, WA
Pinned Tweet
2026 will be the year of tools for agentic thought
2
11
1,325
John Lam retweeted
Made with Claude Opus 5.5. The fearmongering about AI always makes us forget that NOTHING WENT FOOM, as was always predicted. Send this to your doomer friend who has a very high P(Doom). Accelerate.
390
642
4,449
734,571
John Lam retweeted
Your job may be surviving on inertia. Your resentment won’t extend the runway. What do you actually want to do with the intelligence now at your disposal?
Article

You're Already a Meat Proxy

I’m excited about a future that I suspect will be very hard on many long-time friends and certain people that I love. The same technology that is allowing me to be more successful than ever is

141
318
2,392
641,412
John Lam retweeted
The reactor team at @terraformindies has demonstrated successful production of high purity methanol at scale at our Muroc desert test site. Let's review our progress in 2026. In January, we closed on our Muroc test site. By March, construction was underway. By May, we had 1.8 MW of solar installed. In July, we produced CO2 on site (nitter.net/CJHandmer/status/20801…), and in August, hydrogen (nitter.net/CJHandmer/status/20847…). Terraform Industries has built an incredible team to revolutionize not just one or two key technical pieces, but the entire stack required to convert sunlight and air into fuel. Like plants, only 100 times higher productivity per acre and requiring no irrigation, fertilizer, or pesticide. Two years ago, Terraform produced pipeline-grade natural gas for the first time. Since then, we've scaled up our methane reactor, increasing production by 250x and reducing contaminants by 10x, which allows us to increasing revenue per unit of product by 100x because it's now at chemically pure grade. But it was always clear to me that natural gas was only one piece of the puzzle. It is possible to convert natural gas into other hydrocarbon products like gasoline or jet fuel, but given our precursors are CO2 and hydrogen, we can do one better by directly synthesizing methanol as well. Methanol is far easier to convert into longer chain hydrocarbon products, either via the MeOH->DME->ethylene pathway, MTG, or via variations on Fischer-Tropsch. Methanol also has significantly higher revenue per molecule than pipeline natural gas, which forms an essential part of our cost-reduction-through-profitable-scale-up strategy, explained in more detail in our post: terraformindustries.wordpres… TL;DR: If our learning rate is above 11%, our margins will strictly increase as we scale up into a market worth about $10t/year. Not including induced demand. Amir joined the team earlier this year and speed ran the methanol reactor development program, successfully demonstrating recirculation and heat recycling before integrating these new technologies in a production scale reactor, built from the ground up at Terraform's facility in Burbank. As soon as the hydrogen team was done, we shipped our latest reactor up to the desert, plugged it in, and began activating the catalyst. An extended test campaign through the heat of summer followed, punctuated by bunnies nibbling on hoses! The hard work paid off, with Terraform's production of significant batches of high purity methanol, now packaged and ready to ship. A huge congratulation to the entire team, including Amir, Enric, Dominic, Hugo, and @lucie_nurdin. With the completion of this test, Terraform has now demonstrated TRL 9 operation of all of its subcomponents, each developed from scratch, in house, to deliver the highest possible value to the end customer. Terraform will never rest until we have broken the geological and geographical monopoly on oil, unlocking unconditional energy abundance for all of humanity. Join us!
Last week the hydrogen electrolyzer team @terraformindies pulled off yet another first, the sustained production of >99.9% pure H2 from our vertically integrated, California-manufactured electrolyzer stack while it was coupled directly to a solar array at our Muroc desert test site. WTAF? Most commercial electrolyzers need carefully conditioned power from expensive battery-meditated backup systems. Ours runs directly off the sun. Clouds pass, the day turns to night, and we maintain purity from a stack whose Bill of Materials cost is well below $100/kW. Terraform's single-minded focus on capex reduction has allowed us to convert sunlight to hydrogen in an unprecedented demo at a cost below $2/kg. If that wasn't enough, we have a crystal clear plan of steady execution to push that cost below $1/kg in the coming years. Rather than linger on this point in response to experts-with-spreadsheets who said this was not only beyond my team, it was forbidden by known laws of physics, let me tell you a bit about how we actually pulled this off. We started building this test site in March. Once the panels and electrics were in place the CO2 team were the first to demonstrate production on site. Close behind them the electrolyzer team planned the logistics necessary to project substantial operational ability into a hostile test site in the middle of nowhere. It's no good to find you're missing a wrench half way through the day! Much of the test prep was completed before dawn, when the panels go live. The sun came up and the stack immediately started splitting water into hydrogen and oxygen. The team carefully monitored purity and flammability as the sun climbed through the sky. As designed, the stack warmed up and conducted even more power, maintaining solid production until late afternoon when the setting sun shaded the panels. Terraform's synthetic fuel system is uniquely designed to follow the sun and extract the maximum possible value from cheap solar panels. Hydrogen is a pernicious molecule. It leaks through and embrittles metals, burns almost invisibly at a wide range of mixtures in air, burns hot and fast and can easily undergo detonation transition, and has about half a dozen other spookily dangerous properties. My advice is to never work with it unless you absolutely have to. The Terraformer produces and consumes H2 in one compact discrete area with a minimum of complexity and fuss, and as expected this demo was completed in accordance with our rigorous safety standards and no unscheduled excitement! This successful demonstration was also a profound milestone for the team after a testing anomaly last December compelled us to finally rip off the bandaid and move decisively towards the "future design" with half the parts but considerable complexity in assembly. No-one else makes electrolyzers this way and we, more than anyone, know exactly why. And also how to do it anyway, translating directly into a unique cost advantage. A huge congratulation to Ken, Sherman, @ckalitin, Nikhil, Abdullah, and Aaron for their successful test campaign. Terraform's hydrogen and CO2 are the chemical precursors for synthetic methane and methanol, which we make in our own synthetic fuel reactor. Combined, we make oil and gas out of sunlight and air. We are breaking the geological and geographical monopolies on oil production. In the limit, Terraform will deploy these electrolyzers by the millions and they will all be plug-and-play with solar PV arrays. The Muroc smoke test campaign is far from over. We will win!
87
113
1,188
149,362
John Lam retweeted
a quick tip that may surprise some folks an uncached prompt to fable at 500k context window will directly cost you over $5 for A SINGLE REQUEST. that's a cup of coffee or a cheese burger gone. even with subscription quota, this hits like a truck the most common way to fall into that case is when you walk away from a long session and come back after a while when cache expired (claude is 1 hr, codex is 30 mins by default) in particular, when you come back to a long idle session, don't run "/compact" there thinking it'll reduce your cost, because the compaction request is still a real request and it will cost $5 by itself right there the best thing to do is to /compact BEFORE you walk away the next best thing is when you come back and see a large context window, just start a new session, and ask your agent to look for the last session's transcript if it needs context
129
61
1,543
104,843
John Lam retweeted
Astra made a video about their successful Factorio run. Shipping 15,000 ammo to their gun turrets and barrels of oil to flamethrower turrets using construction robots with blueprint item requests makes me laugh. Most people put it on a belt or pipe. piped.video/watch?v=abrWwpGX…
28
27
373
42,943
We have come a long way in 1.5 years!
A new benchmark emerges for agentic AI systems: factorio. You can watch Sonnet 3.5 complete tasks or build entire factories in factorio. github.com/JackHopkins/facto…
1
1
390
Explanation of the setup:
Replying to @stalkermustang
He has both a headless and GUI client available on his Linux VM. The headless supports Lua scripts, which he uses to read game state and plan his actions, but doesn't record replays. The GUI client records the replay and save files, and he takes screenshots and uses "bash('xdotool ...')" or Python equivalent to interact with it. He had some custom MCP servers available for distributing work out to an agent fleet or managing persistent agent sessions, but none of those were used for this task. It could've been done with stock Codex with no custom MCP servers installed.
1
124
John Lam retweeted
GPT-6 Astra has just beaten a standard game of Factorio with enemies enabled after nearly 44 hours in-game time, and 4 days 11 hours on his Codex /goal clock. At API pricing, logs indicate it would've cost about $4500. It could probably be 50x cheaper by offloading more to Luna and providing some standard scripts for common tasks. My next test might be Astra coaching Muse Spark or Qwen 3.8 27B into solving a random seed, now that we have successful logs to review.
165
207
4,054
504,929
John Lam retweeted
astra. 2 prompts. gave it a reference video and the product page. insane lmao react component btw github.com/jal-co/iphone-duo iphone-duo.vercel.app/
20
8
208
49,407
John Lam retweeted
We are excited to announce that AMD and @AnthropicAI are expanding our strategic partnership to accelerate the development and deployment of next-gen AI infrastructure. Tune in at 9:30am PT tomorrow to hear from Dr. @LisaSu from the #AdvancingAI keynote stage! ✅ Up to 2 GW of AMD Instinct MI450 Series GPUs in AMD Helios ✅ AMD has committed to make a strategic equity investment of up to $5B in Anthropic ✅ Deep engineering collaboration across Claude, ROCm and AMD Instinct More on the news: bit.ly/4b2nv5Z
42
191
1,237
202,973
John Lam retweeted
You can now ask Claude about the Anthropic Economic Index, our public dataset measuring how AI is used across the economy. Ask which occupations use AI the most, or what kinds of tasks people are automating, and the answers draw directly from the Index data.
509
918
12,424
2,461,556
The first Vera Rubin clusters are here! Yesterday, @IneffableLabs took delivery of their Vera Rubin NVL72 cluster from @googlecloud @nvidia The AI frontier jumps forward by yet another generation of hardware. Acceleration continues.
165
324
3,436
847,874
John Lam retweeted
This is me talking to my computer without making a sound. After just a month of collecting data, our model is already approaching dictation in accuracy. We were surprised to see that it generalizes to unseen participants as well! (1/n)
238
308
3,753
583,659
John Lam retweeted
Knockoff is now live! Filter out the knockoff crap brands on Amazon. Sorry to brands like WNPETHOME, EHEYCIGA, YXYL, LU&MN, JOYIN, TOMY, GODONLIF, YOOJEE, LINGTENG, LANEIGE, VISCOO, BIODANCE, COOFANDY, BALENNZ, TOSY and LUENX. knockoff.shopping
built a little chrome extension that lets you dim (or hide!) all the crap, mass-produced, fake brands on amazon. should i release it?
1,213
3,904
55,237
11,703,549
i spent a lot of quality time with fable over the weekend building a new kind of "book" (or maybe it's a museum - more later). i noticed that i had to change how i work with fable because: 1/ the "architect" + "worker" model is really good - fable as the architect - i converse only with the architect and it manages the workers that it spawns for different tasks 2/ the architect is responsive while subagents are working - here's a verbatim quote after it dispatched subagents to write some docs. i find that this helps me stay in the loop and not disengage while it is working - this is a really tasteful improvement in fable "Both docs are being written. While they come together, let me confirm your mental model ..." 3/ it spends a lot of effort validating the work of the workers if you work this way. my chat is littered with statements like: "When they land I'll run each through the adversarial refuter, publish the survivors, and add their cards to the landing's walkthroughs section. Next report then." 4/ the level of abstraction i can use is way higher. a throwaway aside in a prompt "here i think we will need to prepare a full vm that contains the exact glibc version, tools, kernel etc. that we want to "freeze and use" for future experiments." and i wound up with a full vm with exactly the tooling that is needed to use it as a reproducible testbed for further experiments. this was sentence 2 out of a 4 sentence prompt - and each sentence was about as ambitious as that one. done in ~40 minutes (that is wall clock time - there were multiple subagents doing things as well). none of this is cheap and i appreciate anthropic subsidizing the tokens for all of us to get a feel for this model. it will be hard to go back knowing what is possible at the new frontier. screenshot below is just the current thread and included 8 hours of quality sleep as well :)
1
2
448
John Lam retweeted
1. as a mental model it is more correct to think of fable+ class models as english -> code interpreters - converts your idea into code into "correct" code regardless of problem complexity and output complexity (diff size). Fable 5 will be the worst of this new class of models 2. diff size/complexity is to be managed purely for review: small diffs - in high risk areas of code (auth/identity/data access/network access/money movement) large diffs for code that can be empirically verified (frontend/backend plumbing/code without network or db access/performance code that can be empirically verified) 3. time it takes to ship software is completely disconnected from time to produce the PR - how long the work takes depends fully on ability to review/merge code while managing risk at scale 4. solving the bottlenecks for above matter enormously- linters/testing/CI/shadow mode verification/empirical verification 5. agency matters enormously- what are the biggest bottlenecks to speeding up the loop and eliminating them? what are the problems that need solving and when do they need solving? what does it take to the solution to all of them today? 6. deep understanding of the full stack matters enormously- what problems are worth pursuing? is there a higher level of problem abstraction to address first? should I give it the sub-sub task, the sub task, or the task itself. what are the major risks with this PR (order of importance: security holes/correctness holes/performance holes). is there a higher speed way of producing data that allows me to merge this? should this be run in shadow or in a sandbox or a flag. understanding every line of logic may not be needed but understanding and managing risk matters enormously. 7. the cost of complexity itself is changing. it might be now worth "maintaining" 50% more code to get a 5% performance win. getting the right abstractions matter less because larger refactors are less tedious. code quality nits become huge drag. very likely, a much smarter model will be maintaining your code so worth taking on more technical debt now. taking the time to hand architect and rebuild systems comes with an enormous cost of velocity 8. if it quacks like a duck and walks like a duck, it's a duck. For low risk cases, it might be more sane to treat code chunks (services / functions) as a black box, like we do for neural networks: do full empirical verification only: has code produced correct outputs for the last 10,100,1000,10k inputs ? can we quarantine this large piece of code - no outbound access to network / database ? what happens when this code is wrong? do we get hacked/or crash(memory/cpu)/is an inconvenience? is it internal facing or external? what can we do to address these risks? 9. eventually, logical verification (line by line review) will come at an enormous cost- save it for where it matters and build systems that are tolerant to empirical verification. is there a decorator that prevents db / network access? correctness bugs are significantly easier to rectify than access bugs 10. what are the rails that allow for even faster iteration? code permissions can be opt in - db writes, db reads, network egress (to where?), PII access. how long does it take to get shadow mode data? how many PRs can be tested? What are the categories of diffs
66
149
1,777
311,664
i've been using fable to build a teams app that runs on @modal, integrates with github, and uses entra authentication. fable 1-shot all of the auth/m365 tenant stuff that nobody can understand and integrated things cleanly with modal. i just had to do the human in the middle stuff with tokens. damn impressive.
1
273
John Lam retweeted
Opus 4.7 has a new tokenizer. This means it's also a new base model. Glory days of pretraining still very much going.
63
126
2,495
328,812
Hey, I'm open-sourcing Clicky. Go forth into the wild and build the future of education and the future of AI interfaces, my friends. I'm happy to have given a spark. Enjoy! github.com/farzaa/clicky
I built this thing called Clicky. It's an AI teacher that lives as a buddy next to your cursor. It can see your screen, talk to you, and even point at stuff, kinda like having a real teacher next to you. I've been using it the past few days to learn Davinci Resolve, 10/10.
282
419
5,909
644,140