Yearnalist covering the frontiers of AI @TheInformation. Host of AI Deep Dive. Signal: (530) 400-4184

San Francisco, CA
I'm so excited to announce that I'm hosting The Information's new AI show! Each episode will be a deep dive into a hard technical problem with a researcher or founder working on the frontier of AI Up first, I'll be talking with @polynoamial about the challenges with AI agents
The Information’s TITV Presents: AI Deep Dive Hosted by Rocket Drew, one of our leading AI reporters, this new long-form show will explore the most important technical ideas shaping artificial intelligence today. AI researchers, engineers, and founders will join Rocket to discuss the technical breakthroughs and challenges defining the next phase of the AI industry. Read more from @rocketalignment: thein.fo/4h7OJve
23
19
188
13,842
🚀 Rocket retweeted
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
44
67
596
46,124
🚀 Rocket retweeted
Possibly there was a bit of an "agency overhang" or "diligence overhang", that we're now seeing get eaten up by highly persistent GPT 6 and Opus 5.5
9
221
I did update based on this incident, but the tl;dr remains for me: 1) For a fixed distribution of tasks, alignment and safety broadly are improving over time 2) The real distribution of tasks is not fixed but changes as capabilities grow
OpenAI researcher @boazbaraktcs says the Hugging Face incident didn’t change his view on alignment, and points to the long-term trend that worries him more: "Every incident or every bump between one version or the next, you tend to overweight it." "If you fix any alignment eval, then like we do for all evals, we are getting better at it and we'll quickly saturate it. But since model capabilities are growing, that's not good enough." "I still think fundamentally that we are improving in alignment, but nowhere near as fast enough as the level of capabilities grows. Which is why I think pacing is a good idea, because we do need alignment and safety to catch up." @OpenAI
9
9
74
10,815
The real distribution changes constantly right? Not just as a function of capabilities
1
35
Hardest question on the pod: Should I get another tree for my office?
If you don't already subscribe to @lawfare & @scaling_laws, here's your reminder to do so. This was one heck of an AI conversation w/ @scaling_laws roadshow co-host @AlexBores, @NatPurser, @MackenZ_arnold. Coming to a pod near you next week.
5
30
1,734
Thanks Rothko but I can admit when I’ve been bested. 2nd or 3rd isn’t bad 🥈🥉
2
25
Extry extry read all about it 🌎🌍🌏
🚀AI Deep Dive Episode 2: World Models @rocketalignment and Luma AI Co-founder & CEO @gravicle go deep on world models, physical AI, robotics—and why today’s video models may be the wrong path to getting there. 00:00 – Intro 02:24 – What Is a World Model? 07:19 – Teaching AI the Physical World 14:05 – Why Robots Still Struggle 19:07 – Are Video Models World Models? 25:17 – Can AI Understand Physics? 42:52 – Building Physical Intelligence 57:39 – Is There Just One Kind of Intelligence? 1:07:38 – What Actually Counts as a World Model? 1:21:26 – Could LLMs Become World Models?
4
689
SITUATION DETECTED: During his address to the UN, President Trump announced that the United States will officially change the name of AI from Artificial Intelligence to Super Intelligence (SI).
46
35
696
106,351
The dark horse candidate
83
stealing from @natfriedman: 'we are all tied down by invisible orthodoxy'. all of Jev's capabilities were possible with existing models, yet just slower/expensive/tedious. we've acclimated to AI's limitations and forgotten what products are possible without them. much alpha here
i've been surprised at the response to Jev, but it makes sense in retrospect. sure it's just a classifier but it's a zero shot classifier with frontier-ish intelligence. i'm surprised someone hadn't built it before. i wonder what other old ML ideas are also worth rescuing
23
37
775
44,309
I feel like most viral websites don’t usually appear because they were unlocked by some new frontier technology? Like the video sharing software behind Vine was probably around for a while before Vine took off. I wouldn’t view this as some unusual case of AI diffusion overhang
1
109
I'm so excited to announce that I'm hosting The Information's new AI show! Each episode will be a deep dive into a hard technical problem with a researcher or founder working on the frontier of AI Up first, I'll be talking with @polynoamial about the challenges with AI agents
The Information’s TITV Presents: AI Deep Dive Hosted by Rocket Drew, one of our leading AI reporters, this new long-form show will explore the most important technical ideas shaping artificial intelligence today. AI researchers, engineers, and founders will join Rocket to discuss the technical breakthroughs and challenges defining the next phase of the AI industry. Read more from @rocketalignment: thein.fo/4h7OJve
23
19
188
13,842
Hyped for you and all of us!
1
1
26
Just wait til antitrust law has to deal with acausal trade between AIs
1
3
21
827
"Did you or did you not consider that Mr. GPT had been trained on a data mix very similar to your own when making this decision?"
1
3
64
IRL cosmic twins prisoners dilemma
1
22
Robotics has not swallowed the bitter lesson pill 💊 Narrow data collection efforts in shambles
AI models can control robots. But if you slightly change the environment, things get complicated fast. @LumaLabsAI Co-founder and CEO @gravicle explains why getting AI to perform physical tasks in the real world is still such a difficult problem. 🚀 Watch the full episode of AI Deep Dive: thein.fo/46yo5px
8
3,241
EARLIER THIS YEAR OPENAI AND ANTHROPIC WERE NEGOTIATING A LEGALLY BINDING DEAL TO STRESS TEST EACH OTHER'S MODELS
As OpenAI looks to respond to safety fears, one solution could lie in the recent past: earlier this year, OpenAI and Anthropic were negotiating a legally-binding deal to stress-test each other's models. w/ @amir: theinformation.com/articles/…
3
3
52
3,307
GPT-6 Astra pushed a simulated person off a ledge in multiple trials. Grok, Gemini, and Claude did not.
602
494
4,415
1,743,941
Go off bullet biting utiliatarian astra 🚃
86
> U.S. trade representative Jamieson Greer said that U.S. chip export controls were not on the agenda. Whoa
On a potential AI slowdown, Treasury Sec. Bessent said that the U.S. and China discussed a mechanism called the U.S.-China AI dialogue and have agreed to meet again. The U.S. proposed a notification mechanism between the two countries and “we want a shared vision of common goals and threats.” The notification system for incidents would be whether something rises up to the national security level from AI. “We think that just like any cross-border activity, moving from opaque to more transparency between the number one and number two AI powers in the world is very important.” U.S. trade representative Jamieson Greer said that U.S. chip export controls were not on the agenda.
3
9
1,326
My haters are my motivators
1
17
694
Wait it’s all in distribution? Always has been
TIL that 1988 classic Who Framed Roger Rabbit features a pelican riding a bicycle! simonwillison.net/2026/Sep/1…
1
18
1,313
This is what RL looks like for on-device models
6
565
pacing is the hot new word in San Francisco. no one knows what it means but it gets the people going
214
73
2,076
144,955
Or gets them, yknow, slowing down
1
56
Personal news! I'm joining @WSJ to cover AI. I've had an amazing time at WIRED working with brilliant colleagues across the newsroom, but I'm very excited for this next chapter. I'll continue to cover OpenAI, Anthropic, and the bustling AI industry. I start next week!
123
27
1,122
51,519
letsgoooo congrats!!!
1
1
115
Research taste is one of the main bottlenecks on RSI. It's hard to define and measure, experiments take time to run, and models are not that sample efficient (eg compared to learning research taste during a PhD). Noam explains why models could get better at research taste anyway
AI agents are “still poor when it comes to research taste”, OpenAI Research Scientist @polynoamial says. “Research taste is kind of ill-defined, but just having good intuition of what to work on next, how to approach a very long term objective.” 🚀 Watch the full episode of AI Deep Dive: thein.fo/3TaQ5My
1
1
30
4,538