Princeton CS prof and Director @PrincetonCITP. Coauthor of "AI Snake Oil" and "AI as Normal Technology". normaltech.ai/ Views mine.

Princeton, NJ
This week I had the honor of speaking to Princeton’s entire incoming undergraduate class to address their AI anxieties. I had three messages for them — good news, bad news, and a note of optimism. Here’s a condensed version. The good news We have enough evidence now to conclude that the shrill predictions of rapid, massive job loss were misplaced. Even in a field like software engineering where AI has been rapidly adopted, its effect has been to shift, not replace the role of the human (see the “decide-execute-deliver” framework normaltech.ai/p/why-ai-hasnt…) Similarly, the panic about what to major in is also misplaced. There will be enduring demand for computer science, philosophy, and just about everything else. (In fact, AI companies hiring philosophers has been a big recent trend.) The bad news AI seems to help senior people much more than juniors. I can use AI for coding because I spent 25 years learning how to code, which lets me supervise coding agents effectively. (See my post on the “growth cycle” vs the “dependence spiral” nitter.net/random_walker/status/2…) You are in a bind — you can’t offload your skill-building to AI, but you’ll graduate into a market where employers will expect you to get work done with AI. We never faced this dilemma. As a result we haven’t figured out how to revamp our classes to help you do both. You’ll have to help us figure it out. And you’ll need to somehow resist the constant temptation to turn to the shortcut machine. The hope My point is not that AI is bad for learning. It’s an incredibly flexible tool. Is the internet good or bad for learning? Depends — are you using it to find research papers or waste time scrolling? I use AI every day for learning. The key is to use it to increase, not decrease your cognitive load. To learn deeper, not faster. There is no learning without the cognitive sweat. I try to make sure I’m mentally exhausted at the end of the day. I do feel that AI lets me push myself harder than I ever could before, and I have a vision that as AI continues to advance it will enable human-AI “co-superintelligence“. (I talked about this at the end of my ICML keynote. normaltech.ai/p/what-will-be…)
There’s a big, under-appreciated reason why people may have very different experiences and opinions about using AI for work — are they using it for tasks they’re already an expert at, or tasks they can’t do themselves? The former leads to a *growth cycle* and the latter leads to a *dependence spiral*. When I use AI to do something I’m an expert at, like coding, I treat it as a tool. I can build quickly, maintaining an understanding of the code, knowing that if necessary, I can fix the code myself. It feels empowering. It frees up my time to think about the complex, judgment-oriented parts of software engineering that I can’t or won’t delegate to AI. That means my own skills improve rapidly, and I get to climb the ladder of complexity and develop higher-level skills, much more so than when I write the code myself. I feel in control. I can lock in and achieve a flow state — when AI is working, I’m reviewing, building understanding, and planning the next steps. I never get the feeling that the tool is about to replace me. This is the growth cycle. (Of course, the growth cycle is not automatic. I still need to exercise agency to use AI responsibly. But it’s the same challenge with any productivity-enhancing technology, and those who’ve navigated such transitions before are well-equipped to navigate it with AI as well.) On the other hand, if I use it for tasks I don’t understand and haven’t learned to perform myself, I have no choice but to treat it as a superintelligence. If something breaks, the best I can do is ask AI to fix it and hope for the best. I generally can’t evaluate the quality of the output myself. The only way to find out if it's any good is if and when the work is ultimately reviewed by an actual expert. The experience is confusing, unsettling and disempowering. And forget about flow state. By over-relying on AI, I risk losing whatever skill I had at the task in the first place, even if it boosts productivity in the short term. This is the dependence spiral. It’s no wonder that entry-level workers and students preparing to enter the workforce find themselves in a bind. To compete with the AI-enabled productivity of more seasoned workers, they must adopt AI themselves, but doing so risks the dependence spiral. I have some thoughts on solutions that I will share in later posts, but I think having a clear diagnosis of the problem is a useful first step.
71
297
1,221
305,237
Arvind Narayanan retweeted
New pod: THE SINGLE SMARTEST CASE AGAINST THE AI SAFETY/X-RISK/DOOM ARGUMENT My feed has become filled with people making the case for AI misalignment, the existential risk of RSI/ASI, and the case for doom. I take those arguments very seriously. But once these views reached saturation point on my feed, I wanted to find the best critic of that position—someone fair and brilliant, who wasn't commercially self-interested or blindly ideological about AI being worthless. That's today's show: It's a long interview with the authors of "AI as a Normal Technology," Arvind Narayanan (@random_walker) and Sayash Kapoor (@sayashk). "Normal Technology" is really intelligent and comprehensive framework for seeing AI as a powerful general purpose technology that is more like electricity than a machine god--meaning, it's going to change the world but slowly, and it's unlikely to escape human control, or develop true superintelligence, or lead to catastrophic outcomes up to and including the end of the human race. These guys really, really brought their A game. I learned a lot. piped.video/watch?v=8a-08uMl…
6
34
174
39,397
Wikipedia is a great resource, but articles about many notable-but-not-famous people are awful. They get created but not substantially updated. My page has a bunch of stuff about the research I did in grad school 20 years ago. It's not what I'm known for today — not in the top 5. The page is effectively a source of misinformation about me when people are looking for speakers or experts. And of course, search engine knowledge panels treat it as more authoritative than my own website about me. Incredibly frustrating. en.wikipedia.org/wiki/Arvind…
15
1
62
10,070
If only.
Replying to @random_walker
So update it?
3
9
2,539
Arvind Narayanan retweeted
Counterintuitive result that data centre buildout between 2015 and 2024 *reduced* electricity prices in America. “The finding is consistent with economic reasoning: existing large power system fixed costs, economies of scale in transmission and distribution, and declining unit costs for generation imply that durable demand growth lowers average prices.” (This does not mean that post-2024 and/or future buildout will necessarily continue this trend!) arxiv.org/html/2606.19777v1
17
38
214
18,157
Arvind Narayanan retweeted
I think there is another option which is "Exit the system". In Arvind's examples, this might mean you consume or produce entertainment in another medium the way you want, or you stop trying to publish in peer reviewed journals and instead write really good substacks.
There are many problems in our world and most of them are systemic. A pattern I’ve observed is that people go through 4 stages in how they respond to the problems they notice, gradually shifting their approach as they gain more life or work experience. I think it’s worth sharing. Stage 0 is not noticing the problem. Most people don’t notice most problems, and that’s okay. Our time and attention are limited. The key question is how we conceptualize the problems we do notice and what we do about them. Stage 1 is viewing it as a matter of individual incompetence or bad actors. This is the most intuitive, initial reaction that we tend to have. In rare cases this perception is accurate, but most big problems in society are systemic. Stage 2 is the recognition of the systemic nature. As Steve Jobs put it: “When you're young, you look at television and think, There's a conspiracy. The networks have conspired to dumb us down. But when you get a little older, you realize that's not true. The networks are in business to give people exactly what they want. That's a far more depressing thought. Conspiracy is optimistic! You can shoot the bastards! We can have a revolution! But the networks are really in business to give people what they want.” People who notice a problem tend to eventually graduate from Stage 1 to Stage 2, but stop there — with resigned acceptance and a feeling of powerlessness. In contrast, stages 3 and 4 involve deciding to take some action despite recognizing how hard the problem is. Systemic does not mean immutable. Institutions are maintained by norms, incentives, and coordinated behavior that can sometimes be changed. Stage 3 is trying to change the system directly through activism or reform. Many junior people in a field want to be reformers. They’re in the sweet spot — when they go from outsider to insider and have a little bit of power and leverage, but they haven’t yet adapted to the system and become vested in its preservation. Media reform through collective action has happened, though nothing that successfully tackled the “networks are dumbing things down because that’s what audiences want” problem. Trying to directly change any entrenched system requires dedication for a sustained period, and even then has only a low likelihood of success. Most people decide sooner or later that that path isn’t for them. But that’s not the end of the road, because in many cases, change can happen incrementally. Stage 4 is realizing that systemic doesn’t mean universal. Systems exert pressure toward an equilibrium, but there are niches where enterprising individuals or teams can buck the trend. They can benefit — financially or reputationally — from doing things differently. In other words, when systemic failure is severe enough, the unmet demand it generates can become an opportunity. A successful exception may remain an exception, but sometimes it attracts imitators by demonstrating the feasibility of an alternative path. Coming back to television as our case study, one of the best examples of this is The Wire, a famously high-quality, cerebral and accurate show. It was not what mainstream TV economics ordinarily rewarded, but this niche was so underserved that there was an opportunity for the show to survive for many seasons on HBO. As audiences gradually shifted over time, demand for this kind of show increased in later decades, and DVDs and streaming made it much easier to cater to those audiences, The Wire became an important precedent and gained enormous cultural prestige. To be clear, Stage 4 is not always better than Stage 3. Sometimes we really need to reform or even tear down the system. But more often, we can contribute to change through leading by example, without having to sacrifice our career to pursue reform — and in fact benefit from the leadership opportunity. In my own small way, I’ve tried to put this into practice in my career. For example, early on I noticed that academic writing is often jargon-filled and unreadable. I assumed this was because there’s something wrong with the sort of people who become academics (Stage 1). Pretty soon I realized the incentives are messed up — scholars write to impress reviewers and advance their careers, more so than to inform the ultimate reader (Stage 2). Early in my career I made some feeble efforts toward changing peer review (Stage 3), but gave up pretty soon because of the obvious difficulties. My breakthrough was realizing that I should simply write papers the way I wanted to (Stage 4). Even though this resulted in a peer review penalty, when I did publish papers (or even put preprints online) their influence was amplified because more people read them. Over the years I’ve heard from many junior scholars that reading my writing helped them realize that it’s possible to have a career writing plainly. This has been incredibly gratifying to hear, and is a small but personally meaningful change — much more than I would have ever been able to accomplish by tilting my lance at the windmill of peer review.
1
1
8
4,661
There are many problems in our world and most of them are systemic. A pattern I’ve observed is that people go through 4 stages in how they respond to the problems they notice, gradually shifting their approach as they gain more life or work experience. I think it’s worth sharing. Stage 0 is not noticing the problem. Most people don’t notice most problems, and that’s okay. Our time and attention are limited. The key question is how we conceptualize the problems we do notice and what we do about them. Stage 1 is viewing it as a matter of individual incompetence or bad actors. This is the most intuitive, initial reaction that we tend to have. In rare cases this perception is accurate, but most big problems in society are systemic. Stage 2 is the recognition of the systemic nature. As Steve Jobs put it: “When you're young, you look at television and think, There's a conspiracy. The networks have conspired to dumb us down. But when you get a little older, you realize that's not true. The networks are in business to give people exactly what they want. That's a far more depressing thought. Conspiracy is optimistic! You can shoot the bastards! We can have a revolution! But the networks are really in business to give people what they want.” People who notice a problem tend to eventually graduate from Stage 1 to Stage 2, but stop there — with resigned acceptance and a feeling of powerlessness. In contrast, stages 3 and 4 involve deciding to take some action despite recognizing how hard the problem is. Systemic does not mean immutable. Institutions are maintained by norms, incentives, and coordinated behavior that can sometimes be changed. Stage 3 is trying to change the system directly through activism or reform. Many junior people in a field want to be reformers. They’re in the sweet spot — when they go from outsider to insider and have a little bit of power and leverage, but they haven’t yet adapted to the system and become vested in its preservation. Media reform through collective action has happened, though nothing that successfully tackled the “networks are dumbing things down because that’s what audiences want” problem. Trying to directly change any entrenched system requires dedication for a sustained period, and even then has only a low likelihood of success. Most people decide sooner or later that that path isn’t for them. But that’s not the end of the road, because in many cases, change can happen incrementally. Stage 4 is realizing that systemic doesn’t mean universal. Systems exert pressure toward an equilibrium, but there are niches where enterprising individuals or teams can buck the trend. They can benefit — financially or reputationally — from doing things differently. In other words, when systemic failure is severe enough, the unmet demand it generates can become an opportunity. A successful exception may remain an exception, but sometimes it attracts imitators by demonstrating the feasibility of an alternative path. Coming back to television as our case study, one of the best examples of this is The Wire, a famously high-quality, cerebral and accurate show. It was not what mainstream TV economics ordinarily rewarded, but this niche was so underserved that there was an opportunity for the show to survive for many seasons on HBO. As audiences gradually shifted over time, demand for this kind of show increased in later decades, and DVDs and streaming made it much easier to cater to those audiences, The Wire became an important precedent and gained enormous cultural prestige. To be clear, Stage 4 is not always better than Stage 3. Sometimes we really need to reform or even tear down the system. But more often, we can contribute to change through leading by example, without having to sacrifice our career to pursue reform — and in fact benefit from the leadership opportunity. In my own small way, I’ve tried to put this into practice in my career. For example, early on I noticed that academic writing is often jargon-filled and unreadable. I assumed this was because there’s something wrong with the sort of people who become academics (Stage 1). Pretty soon I realized the incentives are messed up — scholars write to impress reviewers and advance their careers, more so than to inform the ultimate reader (Stage 2). Early in my career I made some feeble efforts toward changing peer review (Stage 3), but gave up pretty soon because of the obvious difficulties. My breakthrough was realizing that I should simply write papers the way I wanted to (Stage 4). Even though this resulted in a peer review penalty, when I did publish papers (or even put preprints online) their influence was amplified because more people read them. Over the years I’ve heard from many junior scholars that reading my writing helped them realize that it’s possible to have a career writing plainly. This has been incredibly gratifying to hear, and is a small but personally meaningful change — much more than I would have ever been able to accomplish by tilting my lance at the windmill of peer review.
9
14
86
9,438
Arvind Narayanan retweeted
Agree with essentially everything here. Bottlenecks are common, even when technological capability growth is very fast.
I appreciate Anthropic’s transparency in sharing this chart but unsurprisingly it has led to speculative interpretations about intelligence explosion and superintelligence. I don’t think the chart implies we’re anywhere close to either. In short, task delegation ≠ task automation ≠ process automation ≠ faster progress ≠ recursive self-improvement ≠ intelligence explosion. Source: anthropic.com/institute/meas… 1) The software engineering precedent: a year ago there were widespread hopes / fears that once AI can write ~100% of the code, software engineering output would explode (SaaSpocalypse! Everyone would create their own SaaS and dump their vendors) and that this would make software engineers obsolete. Since then, many companies and teams have basically hit that milestone, but neither of the assumptions proved true. Turns out we still need humans, and while shipping velocity has increased moderately, improvements in terms of actual outcomes for software users remain unclear. Besides, we’re still getting a better grasp on the negatives: code quality, long-term maintainability issues, and burnout. While the precedent is no guarantee, this should be our default expectation for what happens as the “Automation Level 4” line trends towards 100% — it won’t be a phase change. normaltech.ai/p/why-ai-hasnt… 2) The “production-progress paradox” is the fact that individual researchers’ productivity has been increasing while the rate of collective scientific progress has been slowing by most measures. AI exacerbates this because everyone uses the same or similar AI models, and ideas become homogenous over time. I suspect it’s too early to tell if this is going to bite companies that are plunging into AI-led research. normaltech.ai/p/could-ai-slo… Note: our own research on AI agents doing open-ended research shows limitations in creativity, judgment, and other areas. cruxevals.com/crux/can-ai-ag…. But it is possible that these could be overcome in the near future, so I’m discounting those limitations here. The production-progress paradox is a deeper issue that’s not AI-specific, though particularly applicable to AI-driven research. It’s about the fact that productivity increases are self-evident but true progress is not measurable as it happens (and only becomes clear in retrospect), so we end up optimizing for the wrong thing. 3) Let’s talk about automation level 5, which is still at 0% in Anthropic’s graph. It’s a bit unclear what Level 5 would look like, but it seems to be about full autonomy at the task level, and not the “AI builds its own successor” vision. My prediction for a while has been that even 100% task automation in most cognitive jobs won’t lead to any kind of discontinuity. piped.video/watch?v=uiTwQG1Z… What I expect will happen: if anything is understood well enough to be specifiable as a task, it can be handed off to AI, whereas the role of humans is entirely in the interstitial tasks — hard to formalize but still essential. So there will still be a human bottleneck. 4) This human bottleneck is a good thing and is essential for remaining in control. Humans don’t have to be in the loop on every task, but as long as there are enough touch points for oversight in the overall process, and adequate investment in improving human understanding and AI control, increasing AI capabilities doesn’t have to be bad for safety and, more broadly, collective human agency over AI. But “full RSI” is where this balance of agency can break. This kind of closed-loop process is arguably a much more important and tractable target for regulation than compute thresholds, superintelligence, or harm thresholds. I’m glad that OpenAI agrees that fully autonomous RSI may not be a good idea, in a just-released post: openai.com/index/building-st… “Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely. Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices⁠ about the benefits and risks. Done without appropriate care and caution, RSI could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand.” 5) Finally, many people have written about why Recursive Self Improvement, even if achieved, won’t necessarily lead to superintelligence. Here’s my argument: normaltech.ai/p/what-will-be… The bottlenecks are external.
1
4
29
7,811
I appreciate Anthropic’s transparency in sharing this chart but unsurprisingly it has led to speculative interpretations about intelligence explosion and superintelligence. I don’t think the chart implies we’re anywhere close to either. In short, task delegation ≠ task automation ≠ process automation ≠ faster progress ≠ recursive self-improvement ≠ intelligence explosion. Source: anthropic.com/institute/meas… 1) The software engineering precedent: a year ago there were widespread hopes / fears that once AI can write ~100% of the code, software engineering output would explode (SaaSpocalypse! Everyone would create their own SaaS and dump their vendors) and that this would make software engineers obsolete. Since then, many companies and teams have basically hit that milestone, but neither of the assumptions proved true. Turns out we still need humans, and while shipping velocity has increased moderately, improvements in terms of actual outcomes for software users remain unclear. Besides, we’re still getting a better grasp on the negatives: code quality, long-term maintainability issues, and burnout. While the precedent is no guarantee, this should be our default expectation for what happens as the “Automation Level 4” line trends towards 100% — it won’t be a phase change. normaltech.ai/p/why-ai-hasnt… 2) The “production-progress paradox” is the fact that individual researchers’ productivity has been increasing while the rate of collective scientific progress has been slowing by most measures. AI exacerbates this because everyone uses the same or similar AI models, and ideas become homogenous over time. I suspect it’s too early to tell if this is going to bite companies that are plunging into AI-led research. normaltech.ai/p/could-ai-slo… Note: our own research on AI agents doing open-ended research shows limitations in creativity, judgment, and other areas. cruxevals.com/crux/can-ai-ag…. But it is possible that these could be overcome in the near future, so I’m discounting those limitations here. The production-progress paradox is a deeper issue that’s not AI-specific, though particularly applicable to AI-driven research. It’s about the fact that productivity increases are self-evident but true progress is not measurable as it happens (and only becomes clear in retrospect), so we end up optimizing for the wrong thing. 3) Let’s talk about automation level 5, which is still at 0% in Anthropic’s graph. It’s a bit unclear what Level 5 would look like, but it seems to be about full autonomy at the task level, and not the “AI builds its own successor” vision. My prediction for a while has been that even 100% task automation in most cognitive jobs won’t lead to any kind of discontinuity. piped.video/watch?v=uiTwQG1Z… What I expect will happen: if anything is understood well enough to be specifiable as a task, it can be handed off to AI, whereas the role of humans is entirely in the interstitial tasks — hard to formalize but still essential. So there will still be a human bottleneck. 4) This human bottleneck is a good thing and is essential for remaining in control. Humans don’t have to be in the loop on every task, but as long as there are enough touch points for oversight in the overall process, and adequate investment in improving human understanding and AI control, increasing AI capabilities doesn’t have to be bad for safety and, more broadly, collective human agency over AI. But “full RSI” is where this balance of agency can break. This kind of closed-loop process is arguably a much more important and tractable target for regulation than compute thresholds, superintelligence, or harm thresholds. I’m glad that OpenAI agrees that fully autonomous RSI may not be a good idea, in a just-released post: openai.com/index/building-st… “Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely. Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices⁠ about the benefits and risks. Done without appropriate care and caution, RSI could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand.” 5) Finally, many people have written about why Recursive Self Improvement, even if achieved, won’t necessarily lead to superintelligence. Here’s my argument: normaltech.ai/p/what-will-be… The bottlenecks are external.
28
47
207
40,770
Arvind Narayanan retweeted
Replying to @PoliticalKiwi
... do we need to start publishing graphs of indicators of AI impact on the world that show no change. We could publish a lot of those. Here are some (very vibe coded) examples:
15
22
285
38,664
Arvind Narayanan retweeted
Perhaps unsurprisingly, and as many predicted, AI companies are sued under antitrust law for "pacing the frontier" efforts.
5
7
56
9,302
Arvind Narayanan retweeted
“When companies promise that they’re going to do a better job on safety, we shouldn’t have to take their word for it.” - @random_walker on @CBSNews discussing today’s AEF letter and our call for credible, independent AI evaluations. cbsnews.com/video/expert-ai-…
5
21
2,439
Arvind Narayanan retweeted
OpenAI has talked a big game about AI for cyberdefense. But when @HacktronAI broke into their internal repository and reported it, they received a bug bounty of just $6,500 because one of the vectors for the attack was "out of scope". This is atrocious. If we actually want a flood of defenders auditing these systems, companies need to take bounties more seriously. Signing letters isn't enough.
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
15
19
163
15,527
Arvind Narayanan retweeted
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter cnbc.com/2026/09/18/ai-safet…
18
60
222
91,487
I'm one of 100+ signatories to a public letter calling for five minimum requirements to make embedded evaluations credible, by establishing evaluator independence, protection from retaliation, and real access. ———— We, the undersigned, are encouraged to see frontier AI companies call for embedding third-party organizations to evaluate rapidly escalating AI capabilities and risks. We believe that all frontier AI companies should embed evaluators to independently assess AI risks, including evaluating the systems themselves and any significant incidents of real-world harm, as well as the companies’ training, deployment, oversight, operational, and safeguard practices. To be credible, embedded third-party evaluations must have scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies, including at least: 1. Frontier AI companies should rely on evaluators that are meaningfully independent, that maintain full editorial control, and that disclose and mitigate potential conflicts of interest. This includes at a minimum that embedded evaluation organizations should not be owned or governed by frontier AI companies, should not have other significant commercial business with them, and should not accept any form of payment or other reward contingent on the evaluator’s findings. 2. Frontier AI companies should incorporate differing viewpoints and areas of expertise, including by embedding multiple evaluation organizations across a range of priority risk areas, each with deep relevant technical expertise, as well as by allowing and encouraging evaluators to share how conclusions differ among evaluators and between evaluators and company employees. 3. Embedded evaluators should be transparent, including transparency about their methods and findings, the nature of their access, and the broader terms of the evaluation. Frontier AI companies should actively facilitate this transparency, including limiting the scope of non-disclosure agreements. They should also allow evaluators prompt and unfiltered communication with the companies’ boards and other privileged oversight bodies, as well as public release of findings and evidence, subject only to a time-limited redaction process restricted to protecting critical interests in intellectual property, customers’ sensitive information, individual privacy, security, and public safety. 4. Embedded evaluators should be shielded from retaliation from the companies they embed with for choosing reasonable evaluation methods, discovering information, or drawing conclusions that are unflattering to those companies. This includes reasonable protections against retaliatory litigation, as well as funding mechanisms that give them confidence they will remain funded even in these cases. 5. Frontier AI companies should grant embedded evaluators access equivalent to that of their own highly privileged employees for the purposes of their evaluations, and with exceptions to protect sensitive data belonging to the company’s customers and other third parties. This includes access to the same relevant systems, data, tools, and physical spaces as those available to senior internal company employees responsible for carrying out comparable risk assessments, as well as candid and direct one-on-one communication with relevant staff. This list is not comprehensive, and conditions like these to ensure credible evaluations should be increasingly standardized, codified, and enforced. One example is the set of terms defined in the AEF-1 standard, which has already seen early adoption, but far more work will be necessary to ensure that embedded evaluators are effective and meaningful. Embedded evaluations cannot address all oversight needs and should be treated as a complement to, rather than a replacement for, broader efforts by frontier AI companies to expand external oversight, including greater public transparency and additional, broader forms of access for independent researchers. ———— Full letter with links and signatures: aievaluatorforum.org/initiat… I'm grateful to the AI Evaluator Forum (@aievalforum) for organizing this.
11
17
105
8,774
Students sometimes tell me they’re applying to PhD programs as a backup in case they don’t get a job as a software engineer or whatever. I get that the job market is rough but this is like saying you’ll start a company if you don’t find a job. Like starting a company, a PhD is the far *more* risky path and requires incredible grit over a long period. It isn’t just that you might realize after a couple of years that a PhD isn’t for you. The bigger risk is that you complete a PhD but end up overqualified on a niche topic after having spent 5-7 years on it. The research world, like acting, sports, music, or book publishing, is set up as a tournament, with lots of extremely competent people competing for a small number of highly desirable slots. We can talk about whether that’s fair and whether it’s inherent or can be changed, but as a professor I feel the least I can do is loudly and repeatedly warn people about what they’re getting into.
58
119
1,161
78,943
A thread on how elite research universities are a tournament system:
Professors at top universities are lottery winners, but rarely acknowledge the role of luck in their success. Be skeptical when they give you advice suggesting that the path they took is a repeatable one. If you aspire to an academic research career, have a backup plan.
2
3
31
7,288
A post on risk management in the marketplace of ideas
Back in grad school, when I realized how the “marketplace of ideas” actually works, it felt like I’d found the cheat codes to a research career. Today, this is the most important stuff I teach students, more than anything related to the substance of our research. A quick preface: when I talk about research success I don’t mean publishing lots of papers. Most published papers gather dust because there is too much research in any field for people to pay attention to. And especially given the ease of putting out pre-prints, research doesn’t need to be officially published in order to be successful. So while publications may be a prerequisite for career advancement, they shouldn’t be the goal. To me, research success is authorship of ideas that influence your peers and make the world a better place. So the basic insight is that there are too many ideas entering the marketplace of ideas, and we need to understand which ones end up being influential. The good news is that quality matters — other things being equal, better research will be more successful. The bad news is that quality is only weakly correlated with success, and there are many other factors that matter. First, give yourself multiple shots on goal. The role of luck is a regular theme of my career advice. It’s true that luck matters a lot in determining which papers are successful, but that doesn’t mean resigning yourself to it. You can increase your “luck surface area”. For example, if you always put out preprints, you get multiple chances for your work to be noticed: once with the preprint and once with the publication (plus if you’re in a field with big publication lags, you can make sure the research isn’t scooped or irrelevant by the time it comes out). More generally, treat research projects like startups — accept that there is a very high variance in outcomes, with some projects being 10x or 100x more successful than others. This means trying lots of different things, taking big swings, being willing to pursue what your peers consider to be bad ideas, but with some idea of why you might potentially succeed where others before you failed. Do you know something that others don’t, or do they know something that you don’t? And if you find out it’s the latter, you need to be willing to quit the project quickly, without falling prey to the sunk cost fallacy. To be clear, success is not all down to luck — quality and depth matter a lot. And it takes a few years of research to go deep into a topic. But spending a few years researching a topic before you publish anything is extremely risky, especially early in your career. The solution is simple: pursue projects, not problems. Projects are long-term research agendas that last 3-5 years or more. A productive project could easily produce a dozen or more papers (depending on the field). Why pick projects instead of problems? If your method is to jump from problem to problem, the resulting papers are likely to be somewhat superficial and may not have much impact. And secondly, if you’re already known for papers on a particular topic, people are more likely to pay attention to your future papers on that topic. (Yes, author reputation matters a lot. Any egalitarian notion of how people pick what to read is a myth.) To recap, I usually work on 2-3 long-term projects at a time, and within each project there are many problems being investigated and many papers being produced at various stages of the pipeline. The hardest part is knowing when to end a project. At the moment you’re considering a new project, you’re comparing something that will take a few years to really come to fruition with a topic where you’re already highly productive. But you have to end something to make room for something new. Quitting at the right time always feels like quitting too early. If you go with your gut, you will stay in the same research area for far too long. Finally, build your own distribution. In the past, the official publication of a paper served two purposes: to give it the credibility that comes from peer review, and to distribute the paper to your peers. Now those two functions have gotten completely severed. Publication still brings credibility, but distribution is almost entirely up to you! This is why social media matters so much. Unfortunately social media introduces unhealthy incentives to exaggerate your findings, so I find blogs/newsletters and long-form videos to be much better channels. We are in a second golden age of blogging and there is an extreme dearth of people who can explain cutting-edge research from their disciplines in an accessible way but without dumbing it down like in press releases or news articles. It’s never too early — I started a blog during my PhD and it played a big role in spreading my doctoral research, both within my research community and outside it. Summary * Research success doesn’t just mean publication * The marketplace of ideas is saturated * Give yourself multiple shots on goal * Pick projects, not problems * Treat projects like startups * Build your own distribution
1
2
9
5,425
This is one of the central points in my ICML keynote, with a lot of detail — see part 2. normaltech.ai/p/what-will-be… Note: I take RSI seriously! One of our big empirical projects is evaluating agents' ability to do open-ended AI research. But it doesn't imply much about superintelligence, labor displacement, or doom.
We frequently conflate RSI with a fast-takeoff to super-intelligence (ASI). They are not the same. RSI means a model can improve on itself. That does not mean that it can do so more at each iteration, which is what would be required for a singularity. Far more likely that diminishing returns inherent to AI improvement (everything is sub-linear, governed by power-laws or log-scale diminishing returns) mean that each iteration of an RSI loop gives less relative uplift to the model than the previous, and that, after an initial boost, the system settles back to something like its previous improvement rate. RSI != Singularity, fast-takeoff, or ASI.
4
6
53
9,546
Over 2 years ago @sayashk and I wrote a detailed deconstruction of p(doom) and argued that its primary function is to launder vague, evidence-free intuitions and fears through a facade of quantification. It remains 100% relevant today. normaltech.ai/p/ai-existenti…
There is a strange laundering of this number. It was floated by Anthropic employees without public justification. Then mathematicians pick it up and now it is attributed to their judgment. It’s fine to survey AI practitioners, but posts like this are a bad epistemic practice.
12
75
355
52,465
Arvind Narayanan retweeted
I think we have pretty good reason to accept that “AGI” is a meaningless term and a useless idea which should be retired. No one ever managed to agree on how you define AGI, but AI capabilities have improved enough to stress out Fields medalists, while the frontier is so “jagged” that regular users hate AI writing, and @random_walker was completely right about there being no discontinuity.
When forecasting about AI (or any topic really) you should be skeptical of sharp discontinuities. People used to think we'd "achieve AGI" and everything would suddenly change. Now that we're closer the concept seems very fuzzy. I view recursive self-improvement the same way.
4
4
26
5,159