unschooled autodidact focused on inner work / self development / spiritual path. technical staff @ si.inc

San Francisco
"I wish there were critiques of the rationalists from sane people with good epistemics" "and where did you learn what good epistemics were?"
4
164
there’s a common pattern where people will believe something from their experience, search for stats or “science” to back it up (however weak), and then pretend that’s the source of the belief. I prefer when people state their beliefs without doing this
Taylor Swift isn't an outlier for going wild over "that's my husband." Wives who rate their husbands as more masculine happier & far less likely to be considering divorce. ✔️ 74% of wives w/ more masculine husbands are "very happy" in their marriage vs. 69% of wives w/ less masculine husbands. ✔️ 64% say divorce is "not at all likely" vs. 45% for less masculine husbands.
6
287
trauma is misgeneralization under the distribution shift of childhood <-> real world send tweet
1
10
328
Ulisse Mini retweeted
I probably watched every single @3blue1brown video from Grant Sanderson (big fan) and I am noticing this for the first time 🤯
151
171
7,018
411,981
lowkey banger
>be me >discover effective altruism >apparently normal charity is inefficient >why donate to random sad thing when spreadsheet can tell you optimal sad thing >fair enough >buy mosquito nets >save lives >numbers look good >feel powerful >couple years later >someone asks an innocent question >why only count people alive today >huh >future people matter too >obviously >my grandchildren shouldn't matter less just because they haven't spawned yet >reasonable.jpg >keep following logic >what about their grandchildren >also yes >what about people in 500 years >sure >5000 years >why not >500 million years >starting to get weird but morality is morality >open calculator >humanity could survive for an astronomically long time >could colonize galaxy >could have trillions upon trillions of descendants >maybe digital people too >maybe simulated civilizations >maybe dyson spheres full of happy uploaded minds >calculator starts smoking >realize currently living humans are rounding error >8 billion people suddenly looking extremely beta >future contains potentially 10^something people >can't even fit beneficiaries in google sheets >new moral priority unlocked >protect the long-term future >stop thinking in units of "people helped" >start thinking in "fraction of cosmic endowment preserved" >malaria? >terrible >but only kills existing humans >AI extinction could delete the entire light cone >nuclear war could permanently derail civilization >bad institutions could lock in terrible values for ten million years >someone invents wrong constitution in 2140 >quadrillions suffer >better fund governance workshop now >friend says maybe we should improve hospitals >explain opportunity cost >friend says hospitals are full of actual sick people >explain scope sensitivity >friend stops inviting me to dinner >need to decide what to fund >easy >expected value >suppose project has one in a million chance of preventing extinction >sounds tiny >but extinction destroys 10^50 future lives >multiply >mother of god >$10 million project has expected value of several galaxies >charity evaluation complete >someone asks where the one-in-a-million number came from >expert judgement >which expert >us >how calibrated >extremely thoughtfully >reduce estimate to one in ten million to be conservative >still beats curing cancer by 38 orders of magnitude >epistemic robustness achieved >someone says maybe project doesn't work >assign 20% chance >still astronomical >maybe project makes problem worse >assign 5% chance >still astronomical >why 5 >because 30 felt pessimistic >publish 46-page report >contains seventeen sensitivity analyses >every sensitivity analysis begins after assuming intervention has positive sign >critic says you're multiplying enormous hypothetical stakes by extremely uncertain probabilities >yes >that's literally why it's important >critic says the uncertainty might be structural rather than numerical >make probability smaller >critic says no, I mean maybe your model is wrong >make probability smaller again >critic begins rubbing temples >discover AI safety >perfect longtermist cause >AI might kill everyone >or create utopia >or seize galaxy >or tile universe with paperclips >or create billions of conscious software minds >finally a problem with numbers big enough for me >start AI safety nonprofit >mission: prevent dangerous AI >hire smartest people available >smartest people immediately start building better AI to understand dangerous AI >interesting >we must understand capabilities to understand safety >we must scale models to study alignment >we must race ahead so less responsible actors don't get there first >we must deploy systems to learn how deployment can go wrong >we must build the thing quickly because building the thing quickly is dangerous >outsider asks why the people most worried about AI apocalypse all work at AI companies >complicated field >company releases stronger model >very concerned >company begins training even stronger model >extremely concerned >company raises $14 billion >concern reaches unprecedented levels >need to influence government >future is at stake >normal democratic process too slow >politicians don't understand exponential curves >public doesn't understand x-risk >experts must guide them >who counts as expert >people who understand x-risk >who understands x-risk >our friends >someone objects that this seems politically convenient >explain we're representing future generations >future generations unavailable for comment >develop concept of value lock-in >terrifying possibility that one ideology controls civilization forever >therefore extremely important that civilization adopts correct values before lock-in >whose values >let's circle back >begin with impartial morality >end with small group of people deciding what quadrillions of hypothetical beings would want >beautiful arc >meanwhile actual humans keep doing annoying things >voting wrong >having parochial attachments >loving family more than strangers >caring about local community >getting upset when told their suffering is cosmically negligible >evolutionary biases everywhere >explain that moral intuition cannot be trusted >except intuition that future digital people count >and intuition that extinction is uniquely bad >and intuition that our probability estimates are sane >and intuition that our institutional choices improve the future >those intuitions survived peer review >someone donates $5k to local homeless shelter >inefficient >could have funded 0.0000000000003% of an AI governance researcher >think of all the simulated people you just killed >okay maybe don't phrase it that way publicly >PR team says "future generations deserve a voice" >much better >journalist asks what longtermism means >say "future people matter" >everyone agrees >great >journalist asks what follows from that >well technically we should redirect enormous resources toward low-probability interventions affecting astronomical futures >journalist raises eyebrow >return to "future people matter" >motte has entered the chat >critic: of course future people matter >me: glad we agree >critic: I don't agree that your institute knows how to help them >me: why do you hate our grandchildren >eventually notice uncomfortable implication >if future value dominates everything >then helping people today mostly matters through effects on future >education matters because future institutions >health matters because future productivity >democracy matters because future trajectory >human beings slowly become instrumental variables in their own moral philosophy >see starving child >feel compassion >check spreadsheet >child's direct welfare contribution negligible >but perhaps childhood nutrition improves national institutional quality >compassion restored >tell myself this is impartial altruism >one day assistant asks obvious question >"how do you know your intervention actually improves the far future?" >silence >open spreadsheet >increase column width >add confidence interval >assistant asks again >"no, I mean how do you know the sign is positive?" >stare into cosmic light cone >10^50 people staring back >none of them exist >none of them can tell me >none of them can falsify my assumptions >realize I have invented the perfect constituency >infinitely important >completely silent >and always represented by me
4
491
Ulisse Mini retweeted
After Jacob Coxon's tweet I've seen thousands of people ask what they can do to make AI killing everyone less likely. My top recommendation: Call your representatives and tell them to prioritize existential risk from AI. We just made callcongress.ai to help with this. We find your local representatives, get their contact information, and provide you with editable scripts to get your point across. Many ways of trying to help reduce risks from AI backfire or require hundreds of hours of effort. Calling your congressperson is one of the few things I think is both easy and very likely to make things better instead of worse.
159
95
491
27,567
Ulisse Mini retweeted
Exceptionally well written statement from an (in my opinion) legendary capabilities researcher at OpenAI. If you are going to read one thing today, I’d recommend reading this one until the end.
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
8
21
303
26,109
excellent statement
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
2
446
Ulisse Mini retweeted
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
523
1,932
8,359
3,082,465
when people focused on X say X is the bottleneck it’s often because they only see the world in terms of X and only see how to improve X
2
13
486
lowkey goated suggestion from the king of cannibals guy
33
1,478
this talk of “preventing human extinction” is obviously a front for the ai companies to increase their customer base by ensuring humanity continues to exist and grow
8
36
620
11,984
Ulisse Mini retweeted
After Jacob Coxon's resignation and extinction warnings, a lot of people are asking 'how could AI possibly kill everyone?' and claiming AI safety researchers have no realistic answer. This is false! Here are the 5 best scenarios I know of: AI 2027: ai-2027.com (I strongly recommend this one for being realistic, engaging, and if you dig into the appendices, highly detailed) Paul Christiano's scenario (Former Head of Safety @ AISI, 2019): lesswrong.com/posts/HBxe6wdj… Gwern Branwen's scenario (widely known independent AI researcher, 2022): lesswrong.com/posts/a5e9arCn… Holden Karnofsky's high-level explanation (RSP Lead @ Anthropic, 2022): cold-takes.com/ai-could-defe… Joshua Clymer's scenario (ex-OpenAI, 2025): lesswrong.com/posts/KFJ2LFog… (There's also the Sable story from ifanyonebuilds.it, though you'll have to buy the book to read that one.) Writing concrete, specific risk scenarios with enormous amounts of detail has been a major research project of many of the most prominent voices in the field! (With the current leaders in effort being ai-2040.com and ai-2027.com)
139
495
2,372
379,954
shame can be performative ime (e.g. shame/attack myself so others don’t shame/attack me) while guilt reflects more what i deeply want but disown
12
349
Ulisse Mini retweeted
"Jacob Coxon [..] got a $20,159 scholarship [in 2022] for the “long term future scholarship program” from the Good Ventures Foundation". AI billionaires: *buys social media platforms, TV channels, streaming services, newspapers for billions* AI bros: *silence*
This post looks like the start of a VERY sophisticated and well-funded PR operation to get support for Democrats to regulate AI into oblivion. Let me show you how it works: 1.) This guy, with minimal followers and no previous account activity, goes to the Wall Street Journal which publishes an exclusive with quotes from him on his resignation 18 minutes BEFORE this post goes up. Planning was clearly done in advance. 2.) Within hours, it has tens of thousands of reposts and the account has 100k+ followers. The post is punchy, quotable, it almost seems professionally written. The first three accounts to quote tweet it all do so within 15 minutes of the initial posting. Remember, this account had basically zero engagement beforehand, so an organic reach explanation seems unlikely. According to Grok those accounts are @_NathanCalvin (General Counsel at Encode AI), @peterwildeford (Head of Policy at the AI Policy Network), and @DKokotajlo (Head of the AI Futures Project), all of which are up-and-coming AI-Doomer policy advocacy nonprofits. The AI Futures Project website says it is funded “primarily” by the Survival and Flourishing Fund, which says on its own website that it has advised Jaan Tallinn, Skype creator and one of the leading investors in Anthropic, to grant over $2.5 million to the AI Futures Project since 2024. Encode AI says on its website that it is ALSO funded by the Survival and Flourishing Fund, which in turn says that it told Anthropic investor Jaan Tallinn to grant $516,000 to Encode AI in 2025. And wouldn’t you know it, the Survival and Flourishing Fund ALSO says it told Jaan Tallinn to grant $2 million to the AI Policy Institute, the 501(c)(3) affiliate of the AI Policy Network, as well. What are the odds that the first three quote tweets of Coxon’s post would all be major AI-restriction policy advocates funded generously by the same donor, who also happens to be one of the leading investors in, and a board member of, Anthropic, the company Coxon was resigning from? And all within 15 minutes of posting (two within ten)? 3.) Jacob Coxon doesn’t have much of a resume, but we do know that, in 2022, he got a $20,159 scholarship for the “long term future scholarship program” from the Good Ventures Foundation, one of the philanthropic vehicles of Dustin Moskovitz, a notorious AI-doomer who has spent tens if not hundreds of millions on policy advocacy to strictly regulate AI, while also being an Anthropic Investor himself. It also just so happens that the 14th person to quote Coxon’s post was @MaxNadeau_ (27 minutes after posting) who is the program officer for the Technical AI Safety team at Coefficient Giving, another of Moskovitz’s philanthropic spending vehicles. Max is not a frequent poster, his last posts before quoting Coxon were before Labor Day, but he was remarkably quick off the mark for this one. 4.) Basically every major Democrat politician and candidate has suddenly glommed on to this post, and conveniently, as the people cry out foe answers, Bernie Sanders already has a bill written to “ban super intelligence” and regulate AI into oblivion, and will be releasing later this week. The bill, among many other things, will create “a new cabinet-level federal agency to safeguard the public from the dangers of artificial intelligence” that will be “advised by an Artificial Intelligence Advisory Board comprised of experts on artificial intelligence.” Do you think, perhaps, Anthropic and its many investors who fund AI policy advocacy might have interest in getting to place a pet “expert” on the board of an entity that dictates what AI is and isn’t allowed to do? And isn’t it fortuitous that this whistleblower came forward with his oh-so scary stories so close in proximity to the release of the most radical piece of AI legislation ever introduced?
Made with AI
7
12
212
5,020
tinder for openai & anthropic researchers who want to quit in sync to avoid accelerating the other side
29
68
1,302
51,387
tinder for openai & anthropic researchers who want to quit in sync to avoid accelerating the other side
4
15
346
48,929
for everyone freaking out about AI near-term, it’s worth listening to this episode with @EpochAIResearch “AGI is still 30 years away” - best counterview i’ve seen & well thought out piped.video/watch?v=WLBsUarv…
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
9
1
40
14,550