šŸ§‘šŸ»ā€šŸ’» growth @seedtalks šŸ“ˆ tracking the AGI foom at itdoeswhatnow.com šŸ—ŗļø mapping british nature at albionguide.com šŸ“± publishing a fun new app winter '26!

East Sussex, England, UK
Guy Parsons retweeted
Um, wow? Opus 5.5: "make the same message much more interesting to a social media audience that loves anime and quick clips and compressed learning" One shot. Also, please do stay for the closing song.
Hey Claude, "Pick a problem or mystery that obsesses you and solve it as best you can & make a movie we can share on social media about it" So it took a crack at the Voynich Manuscript & failed. Then it made this movie, which is pretty interesting to watch and a good explainer.
59
62
1,070
98,135
Guy Parsons retweeted
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’êtreĀ is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seemĀ aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create ā€œhoneypotā€ environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
523
1,935
8,369
3,086,484
Guy Parsons retweeted
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
573
1,909
14,436
1,933,459
it's worth noting that millions (tens of millions?) of us are using agentic AI every day and yet our sessions aren't, by and large, resulting in pervasive unintended cybercrimes
52
14
229
13,900
another one!
We found another cyberattack by internal OpenAI agents, this time targetting @rubygems. They: 1) gained arbitrary remote code execution on rubydoc. 2) developed a novel exploit to steal user API keys (but we do not know if they succeeded). They used package names including hack.rb, evil.rb, inject.rb, and exploit.rb. We thank @j0wimo for initially discovering that agents had posted to RubyGems.
4
2,247
An AI timeline of the past five years: itdoeswhatnow.com/the-story-…
has anyone made a timeline of every important post about ai? like for milestones or posts of cultural significance and public opinion like i should be able to hand the timeline to a friend in some random swedish town and they should comprehend the full ai cycle from 2015 to now
1
1,523
Guy Parsons retweeted
Spiky Superintelligence.
Replying to @GuyP
the frontier radar itdoeswhatnow.com/frontier/ šŸ‘‡
1
1
7
2,957
Guy Parsons retweeted
I asked Claude to make an explainer video of the Hugging Face incident, where AI agents broke out of their sandbox. Made with Fable 5 and Three.js.
We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. openai.com/index/hugging-fac…
5
13
115
25,888
Guy Parsons retweeted
From the OpenAI HuggingFace hack report:
30
76
1,665
76,417
tis not the ai researchers summoning the machine god, but the machine god who summons the ai researchers
Made with AI
2
1,564
what the cynics miss is these AI folks genuinely see the work as more manhattan project-esque than a Big Business. they've splitting the atom, it's a big deal, they have mixed feelings about it!
Replying to @danpfeiffer
I promise you that we are actually literally worried about world-ending consequences from this technology. Please find some people you trust who work at these companies and actually talk to them about their views.
2
3
2,503
'well now what kind of cockamamie slogan is "Now I am become Death, the destroyer of worlds"?! this will never catch on with the public!' YA THINK?
1
1
496
"but why are they doing it then?!" they're scientists. researchers. they can't help themselves! they want to see WHAT HAPPENS. it's the epitome of "so preoccupied with whether they could, they didn't stop to think if they should."
1
292
40
65
1,468
96,173
Guy Parsons retweeted
🪱 AMODEI: ANTHROPIC COULD BE THE LAST COMPANY IN THE WORLD POST AGI Thelonius from WORM TV hitting the streets to see how folks feel about the imminent AGI jobpocalypse
Gavin Baker: Anthropic Believes They Could Be the ONLY Company Left in the World ā€œI would certainly discourage Dario from saying that ever again to anyone.ā€ @GavinSBaker: ā€œInternally, Anthropic is very confident. I have been told by multiple people I trust that Dario has said that Anthropic might be the only private company in the world at some point. Think about that. In this vision, an Anthropic maximalist vision, there's Anthropic, and then there are governments, and that's it. So there's a lot of confidence. I'm sure they have more advanced checkpoints than Fable up their sleeve. They've executed really, really well. I would probably take the under on them being the only private company in the world. Yeah, I think it's gonna be a long time before they land a rocket.ā€ @Jason: ā€œOr they deliver your lunch. I don't know. There might be other businesses in the world. Zipline might want to bring you a burrito. Uber might want to take you somewhere. Waymo might want to drive you home. But sure, Dario.ā€ @DavidSacks: ā€œI might interpret that as a negative signal because it's so hubristic. This is getting into SBF land a little bit.ā€ Gavin: ā€œI would certainly discourage Dario from saying that ever again to anyone.ā€
Made with AI
21
27
276
31,534
Guy Parsons retweeted
The AI safety community constructed a memeplex in which ā€œtaking AGI seriouslyā€ was a prerequisite for being a serious and good person. When inside this memeplex (as many at Anthropic, some at OpenAI, and a few at DeepMind are) your vision narrows until the world feels extremely constrained. The whole future seems to flow through the ā€œone ringā€ of controlling recursive self-improvement. And so even when you worry about AI itself seizing that one ring, you can’t generate better strategies than trying to control it yourself (directly via an AGI company, or indirectly via AGI governance). I’m not saying this is a pure hyperstition. There’s a core truth underlying this perspective: AI will become extremely intelligent and capable, much more than it is today. But the current world is much more spacious and human-empowering than the future which Eliezer originally envisioned (a ā€œbrain in a box in a basementā€ taking over the world by surprise). And it would be even more spacious if this memeplex weren’t active. For example, Satya and Mark and Sundar only started taking AGI seriously because OpenAI forced them to—and even now they don’t really believe in superintelligence—and even if they did they couldn’t get most of their employees on board. Imagine how chill a ā€œraceā€ between Microsoft and Meta and Google would have been, compared with what we have today: Dario and Sam deep in the ā€œone ringā€ memeplex while also personally loathing each other. So the one ring memeplex has an escalating life-cycle. It infects people by letting them harness the narrative that they’re good people for taking AGI seriously, and that making other people take AGI seriously is a boon for the world (despite how terribly that’s gone so far). Then it shuts off their imagination—any sparks of creativity or plans that don’t steer towards the one ring are quickly shut down. Instead they make ChatGPT or the METR graph or other recruiting tools for the memeplex. And yes, they’ll acknowledge that previous versions of the memeplex were too extreme, and led to overly constricted action. But we don’t have time to worry about that, they’ll say, because AGI is coming by 2027/2028, and that’s the end of history. Somehow, though, almost everyone with that view has only a vibes-based definition of AGI. They don’t believe in Dyson spheres by 2028, or self-replicating nanotech by 2028, or brain emulations by 2028. They mostly can’t make concrete predictions, except that it’ll be enough AI that it puts all their plans on a deadline. (Shout-out to @DKokotajlo and @paulfchristiano though, who do make concrete predictions about things going crazy soon.) It seems very hard to break out of this memeplex without just giving up. David Holz is maybe the world champion of that—the only person who was in a position to race for AGI and consciously turned away. Various agent foundations researchers have carved out space to think real thoughts, not the kind of panicky stabbing in the dark that usually passes for safety research. A few others (e.g. Salamon, Hoffman, Vassar, Andre, Sahil, Davidad) are pursuing more unusual paths. And of the people who burned out, I expect some will reorient to doing creative thinking. For others, the main takeaway: yes, the future of AI will be wild. But so far it’s increased peak human agency, and openness to this trend continuing over the next decade will allow you to start creating something worth creating.
the grim thing about the ai boom is everything feels like a distraction outside of the instrumental convergence to RSI
75
135
1,359
274,828
this is about to get drastically weirder in the next 90 days
2
15
3,494