Senior advisor @BrennanCenter | Author 'Dark Mirror' & 'Angler' | ex-Atlantic & Wash Post | Contact: bartongellman.com/contact

USA
Barton Gellman retweeted
Who exactly is spending billions of dollars to try to sway the outcomes of the midterms? New York Times’ Theodore Schleifer, Kareem Crayton, and Dan Weiner join The Briefing with Michael Waldman to discuss what to do about big money and corruption: bit.ly/3VMWGh6
2
1
1,928
Barton Gellman retweeted
What can citizens do to help make sure the 2026 midterm elections are free and fair? Make a plan to vote. Vote early if possible — in person, via drop box, or by mail if necessary. bit.ly/4r1RAZY
1
3
2
1,226
"the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled.Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained"
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
1
3
5
1,615
Barton Gellman retweeted
“Who’s a good boy?!” Hackers just dumped the contents of a Flock camera. They found: 🔴Software explicitly detecting people, not just plates 🔴1.6 million images logged in 21 days 🔴Key to decrypt files stored on the device itself. Finds directly contradict Flock, which claims someone with physical access can't access images. Making it worse,@GainSec warned about the physical access issue more than a year ago & Flock downplayed it. And yeah, the Flock camera logged “Who’s a good boy?!” about every 2 minutes, all while plagued with errors, crashes & reboots. By @dmehro & @josephfcox wired.com/story/hackers-floc…
195
7,588
24,972
879,644
Barton Gellman retweeted
Alright, alright. I bit. My latest on The Antidote: My two cents on recent scary tech predictions. The Future Doesn’t Get Predicted — It Gets Produced carissaveliz.substack.com/p/…
3
54
144
9,860
Barton Gellman retweeted
Today is National Voter Registration Day. Everyone should check their voter registration status, especially if you've moved, changed your name, or haven't voted in a couple of cycles. And if you're a first-time voter, why not register? It takes two minutes.
50
55
2,261
Barton Gellman retweeted
This is what happens when you’re so accustomed to the benefits of modernity - sanitation, clean food & water, low infant mortality - that you forget where it all came from, how much effort it took to build, & how bad things were before. This is the literal definition of decadence
131
2,192
18,540
971,680
Barton Gellman retweeted
BREAKING: OpenAI might have stolen another major proof. In a detailed Mastodon post, which I report in full in the comments, Andreas Thom presents several pieces of evidence suggesting that OpenAI may have trained Astra on conversations in which he and Gábor Kun were working on Gromov’s soficity conjecture, one of the ten problems OpenAI later announced Astra had solved. I know Andreas. We met several times early in our careers. He is an exceptional mathematician, a leading expert on sofic and hyperlinear groups, and one of the most respected scholars in the field. He has spent two decades working on this problem. If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself. And if the allegations raised by Levent Alpöge, Tristan Buckmaster, and now Andreas Thom are all substantiated, we are no longer looking at isolated incidents. We may be looking at one of the greatest intellectual scandals in the history of science. AI is not discovering new mathematics. AI is stealing human discovery.
671
4,835
19,817
1,802,864
Barton Gellman retweeted
Paul is one of the ~3 most credible AI safety experts in the world. Co-invented RLHF, served at US CAISI, a real OG in LessWrong x-risk circles. (His early debates with Eliezer Yudkowsky are famous among the safety crowd and people still cite his "What Failure Looks Like" post from 2019.) He is joining the OpenAI nonprofit board now because he now believes "there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term." Shit is getting real!
46
415
2,740
420,817
Barton Gellman retweeted
To treat an ectopic pregnancy, which is almost never viable, doctors must terminate it. But medical providers are hesitating or flat-out refusing to do that in states where they face criminal penalties for performing an abortion, lawsuits allege. propublica.org/article/ectop…
26
290
539
19,765
Barton Gellman retweeted
Not to brag, but our product has been foldable since 1828.
Apple introduces the foldable iPhone Duo. #AppleEvent
511
15,293
147,258
2,834,025
I love this.
About a year ago, a Princeton student named Kyler Zhou was sitting in traffic in New Jersey. At some point he was so bored that he put on a podcast with… me 🙃 It was my usual talk about the talent crisis. About the smartest kids of a generation going to work for McKinsey and Goldman Sachs, while the biggest problems in the world go unsolved, blablabla, you may have heard it by now ;) What’s interesting is what happened next. As soon as Kyler got home, he emailed me and The School for Moral Ambition (the organization we started to get talented people working on the world's biggest problems). Kyler wanted to bring this to Princeton. Well, we told him no 😅. We had just launched our first university fellowship at Harvard, and there was no capacity for a second campus... Kyler decided to ignore us. Three days into the semester, he launched the first* Moral Ambition Chapter at a university in the United States. In just a few days, more than 120 Princeton students signed up. In November of last year I visited, and the energy was simply AMAZING. Since then, everything has been escalating. Students have started Moral Ambition Chapters at Brown, Stanford, Vanderbilt, Georgetown and 5 other campuses – often before we've heard of them. Which brings me to today. After the launch of our Harvard Fellowship (8% of ALL Harvard juniors applied to the first cohort!), we're now also launching Moral Ambition Fellowships at Princeton, Brown, Vanderbilt and Stanford. The program is simple: twelve super talented students per campus, a $15,000 stipend, and a summer working full-time on the most important causes of our time, like lead poisoning or child poverty or ending factory farming or AI safety etc. – instead of polishing a boring corporate pitch deck. It's just incredible to see how quickly this movement is growing. Every year, thousands of brilliant graduates are funneled into corporate careers that add little to the world. More and more of them are starting to walk the other way. (For students: applications for Harvard, Princeton, Brown and Vanderbilt are open now. Stanford opens October 6. moralambition.org/fellowship…) * Brown University claims that they were actually the first – I've decided to stay neutral in this quarrel 😅
1
7
4,506
Barton Gellman retweeted
🚨 BEWARE of VoteSafe.org 🚨 While it bills itself as a “safe” way to register to vote, it is actually a digital surveillance machine created by Elon Musk’s super PAC.
76
4,186
5,129
70,952
Barton Gellman retweeted
NEW: For months, Katie Miller has denounced ChatGPT, Claude and Gemini in nearly 500 posts. She did not tell her followers that she holds a more than $1 million stake in their top competitor, Elon Musk's xAI. Story w/ @aaronjschaffer wapo.st/4xPWyMd
275
2,238
8,757
1,428,769
SLAPP suits aim to make hard-hitting journalism feel so dangerous and expensive that reporters self-censor. Publishers shouldn’t make it worse. It’s time for them to scrap oppressive indemnification clauses for freelancers, argues FPF’s @SethAStern. freedom.press/issues/oppress…
5
7
3,836
Barton Gellman retweeted
Trying to understand how this story published with my byline last week while I was on my honeymoon. I did not write or review a word of it and it does not pull info from past reporting. Does cleveland.com think they just own my name now? cleveland.com/news/2026/09/c…
35
243
1,520
255,067
Barton Gellman retweeted
The Department of Defense asked OpenAI to provide the U.S. military with a special version of its artificial intelligence technology designed to turn down the military commands as infrequently as possible, according to documents obtained by The Intercept.
12
117
208
27,864
Barton Gellman retweeted
Deploying troops or armed federal agents to election sites is illegal under federal law. And voters have the right to vote free from fear and intimidation. Here is what you can do to protect your vote if federal agents show up at your polling place. bit.ly/4wpSGAH
3
45
68
3,359