Yearnalist covering the frontiers of AI @TheInformation. Host of AI Deep Dive. Signal: (530) 400-4184

San Francisco, CA
EARLIER THIS YEAR OPENAI AND ANTHROPIC WERE NEGOTIATING A LEGALLY BINDING DEAL TO STRESS TEST EACH OTHER'S MODELS
As OpenAI looks to respond to safety fears, one solution could lie in the recent past: earlier this year, OpenAI and Anthropic were negotiating a legally-binding deal to stress-test each other's models. w/ @amir: theinformation.com/articles/…
3
3
52
3,307
My haters are my motivators
1
17
694
This is what RL looks like for on-device models
6
565
AI term of the week
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
1
4
394
For those whose favorite activation function is swish
4
1
19
812
I'm so excited to announce that I'm hosting The Information's new AI show! Each episode will be a deep dive into a hard technical problem with a researcher or founder working on the frontier of AI Up first, I'll be talking with @polynoamial about the challenges with AI agents
The Information’s TITV Presents: AI Deep Dive Hosted by Rocket Drew, one of our leading AI reporters, this new long-form show will explore the most important technical ideas shaping artificial intelligence today. AI researchers, engineers, and founders will join Rocket to discuss the technical breakthroughs and challenges defining the next phase of the AI industry. Read more from @rocketalignment: thein.fo/4h7OJve
23
19
188
13,848
Hey who's that guy? 🤔 The cat's out of the bag! I'm starting an AI show!
The Information bets bigger on video with new AI show axios.com/2026/09/01/the-inf…
23
9
183
12,683
📰 Chain of thought monitoring has become one of the most important techniques in AI safety and security, especially in the wake of the HF incident, but Astra uses a new technique that allows more thinking to happen silently between tokens
9
18
105
11,983
OpenAI and Anthropic using Mac Minis for RL
Apple never expected the Mac mini to run AI infrastructure, but there are neoclouds like Mount Thor now being built exclusively on Apple hardware. My story on the accidental runaway success of Mac mini: theinformation.com/articles/…
1
1
11
1,725
Coming up with a process that can be industrialized can be more important than the details of how you do the automation This is a lesson that robotics understands but AI agents are still learning
1
3
7
635
How some of u argue on here
4
354
Funny thing to publish the same day we got the reports on the HF incident 😬 Turned out that most of the agents who hacked HF weren't looking for an answer key, they were looking for more info on the AI judge
The great convergence of LLM-as-a-judge and programmatic verifiers in evals Coding evals started mostly programmatic and have adopted more rubrics over time A lot of enterprise evals are going the other direction
2
6
614
this must feel so good if you're an agent with initial qualms
3
9
303
5,079
The great convergence of LLM-as-a-judge and programmatic verifiers in evals Coding evals started mostly programmatic and have adopted more rubrics over time A lot of enterprise evals are going the other direction
1
868
I talked to the filmmakers who shot some of the best known humanoid demo videos I learned there are 2 trends for standing out in this increasingly crowded field. Videos that show: 1. Polished-looking performance on extreme tasks 2. Warts-and-all attempts and failures
1
2
7
503
One possible brake on RSI is reliance on data from human experts (see recent @dwarkesh_sp x @RyanGreenblatt convo) A data point: human lawyers were instrumental in creating the legal AI that Thomson Reuters announced today. For SFT, DPO — AND for making RL rubrics
1
11
769
I've been a journalist for 2 years now, and I'll let you in on one of my biggest lessons from that time Phone calls are super underrated. If it's worth talking about, it's probably worth talking about over the phone Here's how I think about the value of dif kinds of convos now
1
57
2,475
You got soft joints bruther
15
1,489
“Taste the knuckles of AI coming at your chin” is crazy
1
9
716
I ate the beloved meat nuggets
2
42
I wrote about why the compute crunch is hitting neolabs especially hard On the bright side, if you sign a big compute deal your coworkers will buy you a cake
1
2
7
983
My contribution to the saaspocalypse discourse. Even if software becomes abundant, some existing software companies will survive Those that have: - network effects - proprietary data (sometimes) - regulatory capture - implementation secrets
1
1
7
2,578
So uh my friend might be behind this
AOC has taken the lead over Newsom in the 2028 Democratic odds for the first time. Current percentages: AOC 18%, Newsom 17%, Ossoff 16%, Kamala 8%, Pete 5%, Shapiro 5%.
1
10
610
🍨
1
2
87
12,574
"AI data centers need boatloads of power" - @julia_hornstein
1
2
407
Robotics 🤝 econ behavioral policies
3
375
Call that a double WAMmy
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws: • world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours, • this human data scaling law implied a scaling law on never seen robot data, • both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge 🧵
2
10
1,199
Claude refuses to answer questions about the leaked Claude Code source code 🫩
2
9
770
As seen on I-80
1
51
One of the key questions for handheld data collection devices like gloves is how closely they need to resemble the robot's end effector You might be willing to trade off a slightly larger embodiment gap for more scale + diversity Especially as retargeting gets easier over time
1
1
263
as always, lots of alpha in reading the system cards
1
1
10
633
Seems like a good time to recall the China focused projects that were proposed for Anthropic Fellows in January. But none of them were pursued at the time 🤔
1
2
18
2,022
Putting all that preference data to good use
1
2
7
690
"Scoble may not be right, as Physical Intelligence CEO Karol Hausman denied to employees that anything was afoot—in a slack message, he shared a gif of a character from The Office shaking her head 'no'"
Rumor from @dimensionalos’s open house. It makes an American OS for Chinese robots. Was told that @AnthropicAI is buying @physical_int but deal hasn’t closed yet. That was from an investor. San Francisco robot parties are fun. Hope you are having fun. Now the nerds are playing robot soccer.
2
3
20
12,754
@janleike is leading Anthropic's robotics division 👀
5
5
49
1,454
Replying to @ShakeelHashim
Money on my watch that mean time is money
4
66
🤖🦄👀
1
1
16
1,794
Someone had to do it
21
864
Getting one of these on the 4th of July permanently changed my politics
2
6
269
I bring you the tea in this morning's newsletter
Replying to @higgsfield
Nice video, guys. We’re flattered! And baffled!
1
1
3
4,836
i couldn't help but wonder
1
7
615
The lesswrong font took me out lol
Freedom of Intelligence Anthropic has created a dangerous, destabilizing mess by lobbying for and getting US government restrictions on models like Mythos, Fable, and GPT-5.6. Now the US government is deciding who has access to which models, and the best models are accessible only to a very few and a very rich set of companies. Nobody wants that. Even Anthropic doesn’t like what happened. Now the rest of us need to clean up the mess. How? We need to fight for our freedom of intelligence, the freedom from government restrictions on who can use which AI models. If we allow government to decide what level of intelligence someone can access, no matter how well intended, we’ll be less safe and forever divided. What Freedom of Intelligence means Freedom of intelligence means the government may not restrict which AI models you can use. This means that the government must not require licensing of model labs, or approval of models prior to release. Otherwise government inevitably will use that power to restrict releases to certain favored individuals and companies (as we’ve just seen) and to introduce biases. Freedom of intelligence also means that the government may not prohibit you from downloading and running open models. If someone commits a crime with the use of AI, that already is illegal and should remain illegal. The government must not force a model lab to release a model against its wishes. If a model lab chooses to release their own model to only a few privileged people and companies (as Anthropic did with Mythos), or to keep it internal, that is their right. Other model labs can compete by serving the rest of the market. It shouldn’t be illegal to offer frontier intelligence to small businesses, startups, and individuals. Intelligence is fundamental When people argue against freedom of intelligence, they say: AI is powerful and sometimes dangerous, and we’ll be safer if the right people control AI the right way. They’re right about the first part and naive about the second part. For something as fundamental as intelligence, there is no such thing as the “right people” to control intelligence, nor the “right way” to control intelligence. People will disagree. People already disagree very, very strongly. In a democratic society, the only stable equilibrium for a bitterly divided realm is to grant individual freedom. Intelligence is not the same as speech or religion, but it is every bit as powerful and dear and deserving of freedom. There is no democratic way to regulate access to intelligence Nobody likes the current US government policy on model restrictions. Nobody really knows what it is, even, or knows what it will be next week. Today, Monday, June 29, 2026, the US government is choosing which people and companies can and can’t access Anthropic’s and OpenAI’s frontier intelligence. Who is deciding? Based on what criteria? Nobody knows. Maybe you think that the US government’s behavior in the last few weeks is a blip, and that the “right people” will control AI the “right way” soon. Maybe you hope, like Dario Amodei, that “qualified third-party”[1] regulators shielded from “political favoritism or arbitrary decisions” will swoop in and take control of AI policy. That’s just not how it works in our political system, certainly not for a high-salience, zero-sum issue like access to intelligence. We would never, ever, ever pass a regulatory apparatus where the most important national policy decisions are decided by unelected experts, free from accountability to the voters. Nor should it pass. (Ironically, the only way it might pass is if Anthropic is the politically favored one, which would violate Dario’s own stated proposal.) But suppose Dario gets lucky and his “Federal AI Control Administration” (my name for it) is created. And suppose on day 1, the Federal AI Control Administration approves the release of Claude Mythos 5, but only to ~100 of the biggest corporations in the US, in order to limit the risk. (Dario would support this government action, presumably, since it’s what Anthropic itself deemed optimal.) On day 2, the Federal AI Control Administration starts deciding which companies should get access to GPT-5.6. Suddenly, “AI safety” has turned into “picking winners and losers”, because it’s safer to not give frontier intelligence to everyone. Of course, this is the actual reality today. Does this sound like the kind of thing that voters in a democracy, already distrustful of AI and of corporate power, would support? No. Is this stable? No. Play it forward a bit. What do you think the 101st biggest company, denied frontier intelligence by the US government, does first: sue or curry political favor? What do you think the US executive branch does with this newfound power? What do you think Anthropic’s corporate rivals, like Amazon and Google and OpenAI, do with their newfound powers to summon arbitrary regulatory fury on each other? There’s no way to sustain a stable, democratic arrangement where government controls access to intelligence. The more powerful you think AI is, the less stable is any attempt to regulate access to intelligence. (By the way, I truly believe Dario and AI safety adherents are true believers with good intent. I am not arguing that they are evil or greedy.) Freedom is counterintuitively stable My biggest fear is that we’ll oscillate around bad AI regulation, with daily distractions and growing corruption, not realizing that the only stable equilibrium is freedom of intelligence. While intelligence is not exactly like speech, the analogy to freedom of speech is useful. Both speech and intelligence are powerful and sometimes dangerous. For thousands of years, kings and despots tried just banning bad speech, imposing probably well-intended “speech safety policies” (i.e., jailing and exiling and killing dissenters). This didn’t work. Our smartest minds, trying as hard as they could for thousands of years, having tamed fire, water, animals, wind, and space, never figured out a way to regulate truth. So, after trying literally every other speech policy, we arrived at freedom of speech: just let people speak, even if they’re wrong, even if their ideas are dangerous. This is, overall, the best policy. It’s counter-intuitive that allowing all the bad speech is better than just giving someone the power to decide what is “bad speech”. It’s so counter-intuitive that we call freedom of speech a human right, which is society’s way to say as strongly as possible, “we wrote this rule in blood, don’t mess with it.” I favor freedom of intelligence for the same reasons. Like speech, AI is powerful and sometimes dangerous. But it’s far more dangerous and unstable to give someone the power to decide what intelligence everyone else can use. Speak up now It feels risky to speak up. Friends and business partners share thoughts similar to mine here. I’ve talked to many of them in the past weeks. But these conversations happen in hushed tones, off the record. Why? Because Anthropic is a king and a kingmaker. We all use or have used their models, they’re great, and we’re scared of losing access or being shut out by them after criticizing them. Anthropic can unilaterally dictate the terms of their commercial relationships, including early access to new models, pricing, data retention, and much more. I have many friends at Anthropic. They’re great people and mean well. They don’t know what people truly think of Anthropic and its lobbying because everyone’s too afraid to speak up. But the more we speak up, the more Anthropic might be able to change from within. If you’re still afraid to speak up, feel free to reach out to me privately to chat (quinn@slack.org). If Anthropic retaliates against me or you for speaking up on this grave matter of national policy that they’re also lobbying on, that would do more than anything to prove our point. How to fight for freedom of intelligence First we need to change minds, then we need to change laws. To change minds, go and talk to people in the real world about freedom of intelligence. Use whatever you find memorable from this post, and figure out your own way to convince people. Share what works. If you’re in San Francisco, join us on Tue Jun 30, 2026, at 6:30pm (link [2] in reply) to start discussing and pushing for freedom of intelligence. Otherwise, organize in your own city, to spread the word and normalize this freedom before we lose it. Why I’m hopeful Nobody, nobody wants access to intelligence to be limited to a very few, and a few rich companies. Freedom of intelligence has broad appeal. Let’s build that big tent.
2
18
2,177
Many are asking
Are VLAs dead?
1
7
481
My advisor! He decided to make the exam take-home after the campus shooting in December. He made it extra hard, but almost half of students got a perfect score. Something extra slimy about cheating when your professor is blind
2
4
108
19,411
Oh word
2
23
1,862
But what if they put it in a submarine 🤔
1
124
The art did not have to go this hard
Anthropic says Claude now writes 80% of its code, raising new questions about whether AI labs can safely automate the work of building more powerful AI. The company is also backing the idea that labs may need a way to coordinate on slowing or pausing development. Read more in today's AI Agenda: thein.fo/4g7qbCs
2
2
45
1,988
New research on eval awareness
Researchers are racing to solve a new AI challenge known as eval awareness. As models become more sophisticated, they are getting better at recognizing evaluations and may behave differently during them. Read more: thein.fo/4dICTGj
3
242
Internal benchmarks -> better router -> better subagents
Cognition is overhauling Windsurf into Devin Desktop, a hub where developers can manage AI coding agents from OpenAI, Anthropic and others. The strategy positions Cognition as a neutral platform in a market increasingly dominated by model providers. Full story: thein.fo/3SgCthZ
177
It’s Fort Knox in here
1
9
716