Chris J. Maddison retweeted
🚨 New paper: Introducing MIND (Monge Inception Distance) Everyone agrees that FID is broken, requires too many samples, slowing down evals. MIND requires 10x fewer samples, is more robust, faster to compute. Our new drop-in replacement for evaluating generative models. 🧵👇
12
68
474
59,625
Chris J. Maddison retweeted
Concentration of power in a few labs is the biggest risk in AI in my opinion
I’m very concerned that during RSI, labs will just stop externally deploying their models. Which means they'll be going full steam ahead on the most dangerous use case of these models (recursive self-improvement), while the public remains in the dark about the nature of capabilities and the state of alignment. And we end up on a path towards tremendous concentration of power.
93
124
1,171
69,175
Chris J. Maddison retweeted
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
1,928
2,858
28,464
7,717,823
Chris J. Maddison retweeted
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
523
1,931
8,353
3,076,864
Chris J. Maddison retweeted
SITUATION DETECTED: Nvidia, Palantir, and Booz Allen Hamilton are restricting or eliminating the use of Anthropic and OpenAI models in some cases due to rising corporate data retention and intellectual property concerns, per The Information.
81
339
4,168
490,017
Chris J. Maddison retweeted
A more serious take on what is happening here. I am in an airport lounge so have some time. Enterprises have made an uneasy truce with frontier labs over last few years: strict contracts that ban training on corporate data in exchange for letting employees use APIs. The problem? Labs don't need to train on your raw data to copy IP.
Exclusive: Palantir, Nvidia and Booz Allen Hamilton are restricting Anthropic’s Fable model for sensitive work over concerns about its data-retention policies. Some customers are demanding irrevocable zero-data-retention guarantees before putting proprietary information into the model. Full story: thein.fo/4dekujW
42
146
1,587
405,410
Chris J. Maddison retweeted
Some thoughts on AI and Theory. 1. To a first approximation, theoretical computer science has been organized around a few major open questions. Much of our work has been motivated by developing approaches to answer these questions. 2. Such “problem-motivated” work has often led to theory-building focused on identifying a general principle that unifies a class of theorems. But much of that theory-building also involved proving new, difficult theorems. 3. Thus, while it’s true that problem-solving was strongly correlated with building understanding, drawing connections, and eventually developing general theories, it would be disingenuous not to admit that our community, perhaps disproportionately in retrospect, focused on and celebrated problem-solving. This was not arbitrary, and was quite defensible. Being able to make progress on central technical questions usually correlated with taste, creativity, persistence, and depth of understanding. Much of our reward structure therefore implicitly relied on the fact that producing an important proof was good evidence that someone possessed these harder-to-observe qualities. 4. It seems likely that we will soon have AI tools available to us that can prove many such theorems in a short time. The cost of obtaining proofs for well-posed mathematical questions will likely fall dramatically. The “scarce” intellectual work will likely shift both upstream: to questions, models and theories, definitions, and conjectures, and downstream: to interpretation, synthesis, explanation, and theory-building. 5. But as long as we believe in humans being meaningfully in charge of our collective decisions and fate, building human understanding of our science (and of science more generally) will remain an essential goal. I plan to expand on this important aspect soon. 6. Historically, finding a solution to an important problem and understanding its significance, implications, and connections were entangled. Finding a proof usually required researchers to discover the right concepts along the way. A dramatic reduction in the time and effort required to prove theorems could break that coupling. We could end up with many more true statements and proofs without a commensurate increase in understanding. Converting an abundance of proofs into human understanding may become one of the central challenges of our field. 7. As a result, I expect the high-level goals of theoretical computer scientists to change. In fact, the advent of powerful theorem provers might help us construct new theories and explore new models far more easily and rapidly, and significantly expand the domains where our models and theories apply. In that sense, the space for theoretical work may significantly expand rather than contract. 8. There’s a high human cost to the disruption that we are likely heading into. Many in our field, and in mathematical communities more broadly, are coming to terms with it. The range of opinions and reactions among mathematicians and theoretical computer scientists is a natural part of this evolution in our thinking as we collectively work through it. Some concrete efforts (including one at @SimonsInstitute) are already underway to think through the immediate scientific and institutional questions arising during this transition.
10
74
298
49,689
Chris J. Maddison retweeted
Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk. The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia. This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.
1,315
1,943
15,232
3,258,400
Chris J. Maddison retweeted
Yesterday's AIs ate the Internet and became a patchwork mediocrity. Today's AI hunts and feasts upon the Mozarts, Einsteins, and Shakespeares of our age. To birth a gang of ghostly geniuses that write our code, drive our cars, do our work, and run our world.
315
575
8,668
422,082
Chris J. Maddison retweeted
The Frontier Lab flywheel is to get the smartest people to use the leading models. Then distill their inputs, your own model’s outputs, and resulting synthetic data and environments. And the people have to use the leading models because they’re locked in mutual competition.
202
210
3,439
270,628
Chris J. Maddison retweeted
Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees. In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn. Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game.
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
84
297
2,505
321,301
Chris J. Maddison retweeted
This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training on user data* with very different privacy/IP implications. Sadly, AI cos don't like to disclose what they're doing. - pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper - use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this - use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs" "De-identification" is weak -- you can identify someone with a small number of bits, and long traces have more than enough. And it doesn't affect IP leakage concerns.
Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
44
154
1,360
281,445
The distinction is not as meaningful as you think.
Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
6
13
448
22,749
AI commodifies every single bit of information it touches. If training on your interactions results in a 0.01% chance that a model surfaces your insight, it's no longer your insight.
2
33
2,680
As leverage goes up, so too does the value of privacy.
my guess, sad as it might sound, is that minor descriptive notes about what tristan & levent were thinking was sufficient to advance the openai team's work to a solution. many breakthroughs can be functionally described in 3 words. 10 bits of information reduces a search space by 3 orders of magnitude.
3
17
6,408
Chris J. Maddison retweeted
Prediction: there will be at least one very large company failure because of shared AI psychosis of its personnel listening and trusting their models a little bit too much.
38
22
687
57,155
Chris J. Maddison retweeted
+1. And while I agree AI may help find "better ways of producing clean energy", we still have to deploy such solutions. So I am not confident it will "stop wars and famine" (which are socio-political/ climate problems).
ENOUGH PESSIMISM IN AI PLEASE I feel that we have become unreasonably pessimistic in our field. 1. I keep hearing AI engineers saying we have to make money quickly because there’s only like 2 years left before we’re automated. Depressing. 2. I see a constant obsession with “having a moat”. This is an incredibly sad mental frame. 3. I keep hearing “we have to catch up”. Soulless. And so on. People: Every solution creates the possibility to attack new real problems. We face gargantuan engineering challenges in our world. How to capture carbon? How to get rid of teflon and plastics in water? How to invent batteries that are at least 30 times more efficient? How to solve clean energy? Better solar cells? Better ways of producing clean energy so we stop wars and famine? How to eradicate hundreds of diseases? Cures for addiction? And so on. Real engineering is about being brave and truly attacking the many problems we face, to engage with a true desire to improve the lives of others and our environment. Good engineering is not about protecting your product to make money at the expense of progress (moat thinking). Good engineering is about ensuring your children and grandchildren will be proud of the choices you made in 30 or 50 years. It is about empowering others. It is about advancing science. It is about being one step ahead. Always, one step ahead, meaningfully, proudly. These are great times. Let’s start thinking positively about all the wonderful things we could achieve together.
4
5
46
9,086
Chris J. Maddison retweeted
ENOUGH PESSIMISM IN AI PLEASE I feel that we have become unreasonably pessimistic in our field. 1. I keep hearing AI engineers saying we have to make money quickly because there’s only like 2 years left before we’re automated. Depressing. 2. I see a constant obsession with “having a moat”. This is an incredibly sad mental frame. 3. I keep hearing “we have to catch up”. Soulless. And so on. People: Every solution creates the possibility to attack new real problems. We face gargantuan engineering challenges in our world. How to capture carbon? How to get rid of teflon and plastics in water? How to invent batteries that are at least 30 times more efficient? How to solve clean energy? Better solar cells? Better ways of producing clean energy so we stop wars and famine? How to eradicate hundreds of diseases? Cures for addiction? And so on. Real engineering is about being brave and truly attacking the many problems we face, to engage with a true desire to improve the lives of others and our environment. Good engineering is not about protecting your product to make money at the expense of progress (moat thinking). Good engineering is about ensuring your children and grandchildren will be proud of the choices you made in 30 or 50 years. It is about empowering others. It is about advancing science. It is about being one step ahead. Always, one step ahead, meaningfully, proudly. These are great times. Let’s start thinking positively about all the wonderful things we could achieve together.
59
160
1,060
113,243
The world is much more complex than the closure of humanity's a priori model. Align yourself with the truth; there is still immense value to be uncovered.
Replying to @avt_im
Everyone who prioritizes wanting to know what is true and why wins from AI.
1
1
4
1,439
Chris J. Maddison retweeted
feels like we're on an increasingly clear path to the full closure of a priori knowledge
5
1
49
14,453