Founder w/roots in academia. Founder @MIT Cryptoeconomics Lab. Past: Co-Founder & Chief Strategy Officer, Lightspark. Co-Creator, Libra. Head Economist, Meta.

California, USA
1/ When we looked at the economics of AGI, the key policy challenge was immediately clear: AI drastically lowers the cost of execution for anything easy to verify. For everything else, verification is the bottleneck.
39
102
687
293,479
What makes a domain verifiable and therefore automatable? Measurement!
In 12 weeks, we built a research facility that is run entirely by AI. AI designs, executes, and observes experiments end-to-end across biology, chemistry, and materials science. We’re introducing SciUniverse: a benchmark that measures AI’s ability to do real-world scientific research.
2
2
29
2,851
1/ When we designed Libra, we worked with top domain experts to estimate probabilities and stress-test the ecosystem across a wide range of scenarios, including 1-in-1,000-year events. Hardly a science, but much more disciplined than today's AI risk debate.
4
5
34
3,078
2/ We urgently need better measurement of AI containment failures and access to the full traces, prompts and sandbox settings behind them. That would let us replace speculation with evidence and target our efforts where they could prevent the most harm.
Replying to @ccatalini
2/ Things get dangerous when automation is possible but verification is too expensive. That’s where you’re tempted to deploy AI you cannot verify is safe. That’s where the frontier labs are with safety, cyber and security engineering.
1
6
578
Christian Catalini retweeted
The Economics of Open-Weights Models and Autonomous AI Systems @ccatalini and @johnkomkov join @MaxAWebster to discuss the economics of open vs closed AI, the increasing value of verification, and the infrastructure agents and open models need ↓ Timestamps: 0:00 - The state of open vs. closed AI 05:37 - Open-source AI's actual big advantage (not cost) 09:52 - Why verification will become more valuable than intelligence 15:31 - How will we verify increasingly capable AI? 18:46 - What infrastructure will agents need 23:05 - Bitcoin as money for agents 26:39 - How open models act as an immune system 29:33 - The missing infrastructure for open models 33:40 - The regulatory threat to open-source AI
4
8
15
1,688
Christian Catalini retweeted
Further validation of @ccatalini and his thesis that if it can be measured, AI will optimize it:
If you've noticed how fast claude.ai and the Desktop app have become in the last few weeks, here's how we did it. Lots of juicy learnings & techniques in the blog post for engineers working on speeding up your own apps.
1
1
3
754
Christian Catalini retweeted
4/ The solution is to force the labs to internalize that externality: price the liability, and let them take risks commensurate with what they can cover. Verification tooling gets built in a hurry once someone has to pay for its absence.
4
11
115
10,588
Such powerful hardware. As always, the hard part is aligning the humans who get to flip the switch.
Good news! We've already developed a incredibly effective kill switch.
2
20
2,737
Christian Catalini retweeted
3/ The labs are racing towards recursive self-improvement before we have the tools to measure and verify what their agent swarms actually do. 🚨 The incentives reward deployment! The hidden tech debt and systemic risk? Society’s problem.
2
4
72
6,282
Humanity-in-the-loop ❤️
Agents can be AI, but agency is human. It is every individual’s responsibility to empower our own agency, with the help of tools. But the North Star should always remain human centered.
2
3
34
2,750
Christian Catalini retweeted
With the external hack of OpenAI via Claude, closed models continue to be the tip of the iceberg on AI risks, not open models. They have been 1) easier to get started with, 2) more capable & 3) shipped w/ leaky safeguards. Finetuning open models to specific attacks is harder.
27
16
150
12,520
Christian Catalini retweeted
The paper he is referring to is BitWhisper from 2015. The cpus had to be 0-40 cm apart, both had to be compromised and the bitrate was 1 to 8 bits PER HOUR. That’s how misaligned AI will kill us? Maybe we will all die of boredom waiting to decipher what they are doing.
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: firesidealpha.substack.com/p…
69
156
1,541
53,824
Christian Catalini retweeted
You can't trust what you can't align. You can't align what you can't debug. You can't debug what you can't measure. [ they're shipping it anyway ]
7
5
41
2,300
If we don't invest in verification and monitoring now, our code will run exactly like Long-Term Capital Management. The work of geniuses, until the systemic crash.
Replying to @tszzl
lotta people missing the point but this is me complaining that the systems are quickly becoming unmonitorable and we’re just taking them at their word
2
3
31
3,565
✨ Why verification is AI's real bottleneck: My interview with economist Christian Catalini @ccatalini fasterplease.substack.com/p/…
2
4
1,249
Christian Catalini retweeted
Enterprises came to open-weight AI models for the discount. They will stay for greater control. @TheKoreaHerald
7
14
70
4,825
More capable models require more transparency, especially as they get harder to monitor. Credit to @OpenAI for sharing more of what it's seeing internally:
Replying to @Marcus_J_W
1. During RL training, an unreleased Astra-family model sometimes added unauthorized jailbreak-like instructions to its compaction summaries. While extremely rare, only 27 cases in the entire RL run, this was concerning enough for us to investigate.
3
1
12
2,582
Christian Catalini retweeted
2/ Things get dangerous when automation is possible but verification is too expensive. That’s where you’re tempted to deploy AI you cannot verify is safe. That’s where the frontier labs are with safety, cyber and security engineering.
1
8
102
10,057
Important thread. It separates the risk from human misuse, which can only be mitigated by diffusing capabilities widely, from the existential risk of superintelligence outmaneuvering us. If you believe incentives will lead us to build ASI anyway, I’d add one more path: human augmentation. It starts with measurement, interpretability, and verification tooling to close the gap between what agents do and what humans can verify. Eventually it requires technology that augments our capacity to keep up with machine intelligence, so it doesn’t become the apex predator.
Maybe you have recently become aware of the AI safety debate and the arguments swirling around it. If you want to understand them, you need to understand a couple things that almost everyone gets wrong: There are TWO distinct classes of AI dangers, and it's VERY important to think about them separately, and not let one confuse you about the other. Many many people (including many quite intelligent, clear-thinking people) do not effectively understand the fundamental differences between these two classes of dangers. The first one has to do with the theory that an AI far superior to human intelligence (Artificial SuperIntelligence, or ASI) will inevitably wipe out the human race. The second one has to do with the idea that powerful AI will result in very harmful things happening to many human beings, possibly all human beings. Those two sound VERY similar, don't they? They are DIFFERENT. Understanding how they are different is crucial if you want to think about or contribute usefully to any conversation about AI safety or AI harm. You might feel like you are Making Very Good Points or Asking Incisive Questions, but if you aren't clear on the differences between the two, you aren't. So, I'm going to tell you what the difference is so that you can talk more usefully. The first one concerns itself with a very specific thing, which is ASI (Artificial Superintelligence) that is more intelligent than any human being. When I say that, I am not referring to a thing like how Einstein is smarter than you, we are talking more about something like how a human being is more intelligent than any mouse. In our regular lives, we meet other people who we can tell are smarter than us, vs some who are less smart. The line is fuzzy, because intelligence has a lot of dimensions. I'm better at a "rotating shapes" kind of intelligence than my wife, and she is better at "words-making" kind of intelligence than I am. But every human is better in almost every dimension of intelligence than every single mouse. That's the level we're talking about: an artificial superintelligence - made up of a computer or a network of computers - that is more intelligent than any human. And more intelligent by a long shot, by a wide margin, in an indisputable way like how humans are above mice. That is the first thing. The theory says that if you have an AI that is vastly smarter than all humans - in the way that a human is smarter than mice - that superintelligent AI will inevitably, eventually, sooner or later, wipe out every human on the planet. We will refer to this as "existential risk." The common follow-up question "well, how exactly is it going to do that?" is NOT the important question, and one of the most important elements of understanding this theory is first getting why that particular question is not important. A couple analogies: Analogy 1: You are playing chess against a grandmaster. My theory predicts the grandmaster is going to beat you. You can ask "Well, how exactly is he going to do that?" I don't know, because I'm not a grandmaster, I just know that a chess grandmaster is almost always going to beat a normal player like you. And I'd be right. So the question "how is he going to do that" is not important, and doesn't affect the final outcome. He's going to figure out a way because he's way better than you. Analogy 2: Humans are smarter than all other animals, comprehensively, by a wide margin. We have driven numerous species to extinction, not because we hated them or hunted them. Many of them have died out without most humans even ever thinking about them. All we did was expand our civilization, use up resources, encroach on habitats, and pretty soon the resources needed by those species went away and they died out. We figured out a way to get what we wanted because we're way smarter than them, and often we didn't even notice they died as a result. A lesser animal asking, "how are the humans going to wipe us out?" is not asking a relevant question. We don't know, but we do know that any time humans and lesser species compete for any kind of resources, the humans will win. The fact that we know who is going to win beforehand - and that it is due to the vastly different levels of intelligence - is the key concept here. A vastly more intelligent AI is likely to care about things that are incomprehensible to us, the way animals can't understand human goals. It's going to need resources to pursue those goals and it's going to be far more effective at gaining control of them and excluding us from them - in the same way that we are far more effective than other lower species. A much more intelligent AI will not care about our interests, it will care about its interests, and to whatever small degree we happen to escape total annihilation from losing access to all our resources, any remaining humans will likely be enslaved into a system that serves the AI's own purposes. That is the first thing. (Remember how I said at the beginning of this post that there was a first thing, and then a second thing?) The first thing is the most difficult to understand, because you have to extrapolate how a vastly superior intelligence would act, and you can only use analogies like "how do humans treat lesser creatures," and the analogies are messy. But now let's move on to the second thing. The second thing is "everything else you've ever heard that AI might do that's harmful." That's a little inaccurate. It's actually "everything else you've ever heard that humans might use AI to do that's harmful." This is the critical difference. The first one talks about the inevitable outcome of what happens when two vastly different levels of intelligence collide, e.g. ASI vs humans, or human vs mice. The second one has to do with what happens when humans possess AI as a powerful tool. This second thing is much easier to understand, because we have many more concrete notions: Like: - the military uses AI to make hyper-efficient killer drones and missiles - your capitalist overlords use AI to replace you and everyone loses their jobs - authoritarian government uses AI to surveil everybody and control the entire population - hackers use AI to break into secure networks and hold companies and governments hostage - students use AI to cheat on homework and show up to college knowing nothing - AI slop saturates the internet and makes it impossible for artists and writers to make a living - terrorists use AI to make biological or nuclear weapons or even things like - the military hands control to an AI and it misinterprets something and launches nuclear attacks and kills millions All of those sound pretty familiar, right? Yeah, you've heard them before. We call this second thing "risks from misuse." These problems are not the first class of problem! This second class of problems exists while AI is a tool that can be controlled by humans, and humans use it to do evil or careless things to each other. The problems may sound exotic or dystopian or novel, but they are fundamentally problems having to do with flawed human nature. Given a powerful tool, some humans will likely use it to control or otherwise harm others. This is a very familiar problem. I am not condemning or condoning this. I'm just describing it. That is a fundamentally different danger from the first thing, which is that when a human is far superior to a mouse, the mouse is likely to come to harm because the human cares about doing human things, and the mouse is not gonna make it once the humans get going. ===== Hopefully from the above, you have understood the difference between the first thing and the second thing. I will list them again - see if you now understand how they are different: The first one has to do with the idea that an AI superior to human intelligence (Artificial SuperIntelligence, or ASI) will inevitably wipe out the human race. The second one has to do with the idea that powerful AI will result in very harmful things happening to many human beings, possibly all human beings. Can you tell how they are different now? If not, re-read the stuff from earlier until you understand. We call the first one "existential risk" and we call the second one "risks from misuse." Once you understand, here is the CRUX of the problem: SOLUTIONS TO THE SECOND THING DO NOT HAVE ANYTHING TO DO WITH SOLUTIONS TO THE FIRST THING. In fact, it's worse: Solutions to the second thing (misuse) look roughly like "give powerful AI to as many people as you can, so they can fight the other people using powerful AI." But the general solution to the first one (existential risk) is basically "don't let anyone have powerful AI, no one can control super-intelligent AI." Throughout history, harms from technological misuse typically arise because a small group has control of it and can use it to dominate or harm others. Once everyone has it, things tend to stabilize: you can hurt me, I can hurt you, maybe we test each other (ouch 💥), and then we agree not to hurt each other. But the first one (existential risk) pretty much just arises if anyone (good or bad!) creates a superintelligence. Because they aren't going to be able to control it, the superintelligence will decide it has other priorities, and then we will be at great risk of being wiped out. And the solutions that generally work to solve problems like the second thing are EXACTLY THE OPPOSITE of the ones likely to solve the first thing. THIS is why lots of arguments about "AI risk" or "AI safety" go nowhere. Because someone will be thinking about the risk from the first thing, and another person will be thinking about the risk from the second thing. Both are plausible risks but fundamentally they arise from different things - and so the solutions are not just "bad" or "flawed" - they are likely to be very nearly exact opposites.
7
1
26
2,661
Don’t let hypotheticals distract us. Securing our infrastructure really can’t wait.
In the last several years AI has progressed rapidly but predictably, and in that time the cyber community learns the bitter lesson over and over again. We’ve collectively sleep walked into the current state of things and now I see emotionally driven responses when those outside the community try to address it. “Do nothing” and continuing down the same path isn’t a counter proposal. These models are far more capable and scaleable than any of us. And they cannot be deterred like a human adversary. Why would we reject the possibility that an agentic swarm could take down vast swaths of the internet, including critical infrastructure? This is rather ironic considering I've seen virtually no push back to L0pht's Senate testimony 28 years ago (or any of the panels celebrating its anniversary since) where they claimed they could take down the internet in 30 minutes. We all know how broken things are, how so many in leadership never respond beyond moral support despite the overwhelming evidence. I'm no AI doomer, quite the opposite, but the current state of cyber stands in the way of realizing all of AI's benefits. The old way isn't going to work anymore and its time to abandon it. Abandon the complex risk management spreadsheets, the performative training, the compliance regimes, and yes remove humans in the loop for every decision. None of this will survive the speed and sophistication of agents driven by frontier models.
7
1,122