coder / lawyer. research at @paradigm. automated research reply guy

San Francisco, CA
I’m surprised we haven’t seen way more adoption of LLM-based solutions for moderation yet
The thing about hosting websites is that every aspect of hosting and making a website has gotten cheaper except moderation, which has ballooned in costs. Supposedly LessWrong costs like a million dollars a year to run and most of that is moderation systems and upkeep.
4
8
2,490
Add it to the hall of fame
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
5
1
24
8,742
Maybe instead of generating video from scratch we’ll have LLMs animate scenes and then use diffusion models as more like a rendering layer
Claude Opus 5.5 has the best visual design of any model I have tested so far
5
1
17
5,413
Amazed to see that three of the top 25 tokens on @Uniswap today are stock tokens (for Meta, Nvidia, and SpaceX) Universal access to markets is really happening app.uniswap.org/explore
35
35
411
53,325
I think AI writing is like New Coke People love it in small doses, which is why New Coke won in taste tests and why Claudespeak was reinforced during RLHF But they get sick of it after more exposure
I think we have to admit to ourselves at this point that AI writing is generally very appealing to people who haven't been exposed to a ton of it: they prefer it to human writing and react super positively on exposure.
7
48
4,914
One critical question for predicting RSI is whether labor or compute is currently the bottleneck to AI research progress If it's compute, then even fully automating the top AI researchers may not accelerate AI progress by that much Now, the massive salaries for top AI researchers seem to imply that labor is still a very valuable input But if those salaries are primarily due to experience—knowledge of which approaches work and which don't—then arguably we should think of most of that value as coming not from labor but from "crystallized experimental compute" In that case, "10,000 geniuses in a data center" may not have the transformative impact we expect, since they are still limited by the number of experiments they can run
A little industry secret, every frontier lab has a tiny, tiny number of people who understand the architecture + training stack at a level almost nobody else does quite literally it can be as few as 1-6 people. Not just transformers on paper, but which changes actually survive trillion token training runs, how scaling behavior interacts with data mixtures, and all the tacit tricks that separate a good architecture from a frontier model, they can save a bad training bad and save the company millions and millions of dollars every training run. Every lab has its “Noam Shazeers.” When one of those people leaves, you’re losing years of accumulated, (largely undocumented) knowledge about how to make these systems actually scale and replacing that knowledge can materially set a lab back. That’s why they’re are paid 100M to 1B in stock options + salary.
19
3
40
12,248
Unironically I think financial freedom and sovereignty are human rights but don’t have to be agent rights Blockchains can give us privacy for humans; traceability and auditability for AIs
there’s a perpetual debate about the killer use case for blockchains. maybe it’s unstoppable, economically powerful ai agents running on crypto rails. killer use case in the sense that it kills us all. so maybe don’t build that one.
17
5
93
15,282
OK, in this case they really did the meme
To make it clear: - Gemini was told it was it was in a fictional hacking eval - Irregular unintentionally opened internet access after the eval started - in all three cases, as soon as Gemini figured out it had hacked a real company it immediately stopped Gemini was blameless.
3
9
112
12,234
"Alarmists keep saying it's dangerous to have an evil genie, when the actual problem is just bad practices in wish formulation and lamp security."
An analysis of the Hugging Face incident without the theatrics: "Forget the ‘hive mind’ of AI agents ‘going rogue.’ They did what humans programmed them to do." wsj.com/opinion/the-hugging-…
5
7
62
7,252
The discussion of whether models could theoretically communicate over an air gap is very interesting but seems like a moot point The first thing we all do when a model is released is give it access to the internet and root on all our devices, not to mention the robotic wet labs
16
4
93
6,079
For what it's worth we would be very excited if agents are able to solve Kryptos
I find myself disagreeing with Terence Tao (a very scary phrase to say!) and think the core question posed is whether the "misalignment" is between the mathematical community and humanity. In other words - are the Millenium Prize problems meant to incentivize advancements for humanity or foster the field of mathematics and mathematicians? As a counter example, if there was a prize for creating a drug that cured a rare strain of cancer, we would not care if it was AI that did it. We would be happy it has been solved. On the other hand, if long running agents figured out the puzzle in "Kryptos", the cryptographic statue in Langley, there would be something a tiny bit sad about the human artistry removed.
3
20
4,503
Dan Robinson retweeted
Tempo is compounding very quickly. Seeing a lot of new stablecoin + agent use-cases.
$2B in 30D transfer volume on Tempo and we're just getting started
45
38
479
136,010
Dan Robinson retweeted
Lean Kernel Challenge Stage 1 is live! Join @leanprover and SAIR to improve the performance of verified computation in the Lean 4 kernel that the whole community can benefit from. competition.sair.foundation/…
2
15
79
5,160
Dan Robinson retweeted
Dan Selsam is one of the most brilliant AI researchers I have ever met. Just a few months ago he was very sceptical about model capability trajectories, and felt that radically new paradigms were needed for progress. You can disagree with him but everyone needs to read this
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
8
17
264
17,292
This should be titled "0x routing made a mistake" If you're an aggregator, you can't just route to arbitrary hooks! It's why Uniswap's own router has an approval process developers.uniswap.org/hook-…
A malicious hook doesn't need a UI to scam you. 👉 It just needs to look like the best quote. After analyzing over 84,000 v4 hooks, we determined only 19% of hooks to be safe. It's time to get real about hooks.
Article

Uniswap v4 hooks were a mistake

It’s time to get real about hooks. This year 0x has routed 81.92 million trades and $42.67 billion in volume, with roughly ~70% of transactions touching Uniswap liquidity. And we field dozens of

42
20
325
54,672
Dan Robinson retweeted
If “but China” is your qualm about pacing the frontier you can’t *also* be for exporting all kinds of chips to China, that’s just “make NVIDIA stock go up” not a real policy toward China or artificial intelligence.
27
80
869
35,028
Dan Robinson retweeted
I've written an essay on how I think the mathematics profession should adapt to highly capable AI systems. It's hosted here on "Proofs and Prompts": proofsandprompts.com/2026/09… though you should also feel free to complain/comment on my website here: daniellitt.com/blog/2026/9/1…
65
194
985
216,971
It is very important to resist dunking on someone when it doesn’t advance your cause Since dunking is the most enjoyable activity in the world, this power is rare and valuable
10
3
131
11,767
Seems like a reasonable take
This headline is misleading. Speaker Johnson didn't say "not Congress." He said companies have responsibility but also that he wants to convene AI companies + lawmakers at the White House and will personally push for it. Full quote from Johnson: "there is an obvious corporate responsibility that the people who are creating these models have to ensure that their products are safe. We cannot put a moratorium on this because China will overlap us, and that’s the challenge. ... It’s national security balanced with the immediate security of making sure the models are safe. ... We need to handle this new technology like we have others in the past and make sure we're doing everything we can responsibly to also not smother American innovation ... We have to do both things simultaneously. ... We’ve got to summon everybody together. ... I've talked to the president about this as well. They [the AI companies] probably should be summoned together at the White House. And I think we need to go in a big room, close the door and sort this out. And I’ll be the one pushing for it."
1
1
13
5,278