trying desperately to make ai love us at coefficient giving

Anthropic's implied valuation in secondary markets fell around 5% immediately after Dario published We Must Pace the Frontier (onchaintimes.com/pre-ipo-per…). This seems like important evidence (albeit weak, secondary markets are illiquid, weird and opaque) that the market does not think Pacing the Frontier is in Dario's financial interests. I think it is clear that Dario/Sam/Elon are not financially incentivised to call to slow down development. They are calling for slower progress because they are genuinely scared of AI's risks, not because of regulatory capture. Interestingly, I don't think Andreesen and Sacks have a clear financial interest in their position either*. If the frontier gets regulated, AI could become more commoditised, which might be good news for the application layer companies and neolabs that form such a large part of a16z's portfolio. Both sides of this debate seem ideologically motivated rather than financially motivated to me. People should spend more time evaluating AI doom arguments and less time scrutinising financial incentives. * OTOH I think Jensen does have a clear financial incentive to be anti slowing down, and this does seem important to me to highlight.
Replying to @sriramk @micsolana
I mostly agree (mostly you should engage with people's arguments not their incentives) but I think it's not completely symmetric. There's a long history of industries downplaying risks to suit their financial interests (tobacco, oil companies) and no examples that I know of other than AI of industries overstating risks against their financial interests. To put it differently, it's much more established that financial interests can corrupt people's thinking/statements than other interests. Obviously, people are claiming that Dario/Sam/Elon are financially motivated (they're trying for regulatory capture) but I think the view that their position would enrich them is much weaker than the case that Jensen's position would enrich him, which is where the asymmetry lies. In fact, since they came out in favour of pacing the frontier, I have heard that the valuations of Anthropic and OpenAI have fallen substantially in the secondary markets!
4
25
248
22,920
Jake Mendel retweeted
Amodei, Altman, Hassabis, and Musk have spoken (albeit somewhat inconsistently) about the need to slow down to avoid AI existential risk. Zuckerberg, Huang, and almost every VC say it's bullshit. The difference is not that the first group is conflicted but the second isn't. They both are super conflicted. If AI existential risk was to hype up the technology, why would Zuckerberg, Huang, and VCs be dismissing it? If it were "regulatory capture," why would Hassabis and Musk, neither of whom are at the frontier (sorry), be discussing it? The difference is that the first group has been thinking about these risks seriously (at least off-and-on) for over a decade, since well before they ran large AI companies. The second has very clearly only engaged superficially, and is now annoyed that the first group's actions could dampen the massive profits of the AI boom for risks that sound like science fiction to them. The terrifying reality is that these concerns are not being made up to sell a product or a policy. They have been around in some form since Turing, and fleshed out in great detail before OpenAI or Anthropic were a twinkle in their CEOs' eyes. They are fundamental to a technology whose entire purpose is to surpass human capacity, unless you are able to solve the very, very difficult problem of aligning the interests/goals of a more powerful agent with something as complex as human flourishing -- which the companies are very visibly failing to do. They need to simply not build superintelligence, using anything remotely like current techniques, based on anything remotely like the present understanding of AI.
20
82
654
25,220
Jake Mendel retweeted
Really happy to see @polynoamial thinking about how to ensure alignment during RSI, this is an extremely important question! The way I see it, there are (at least) three problems with using alignment evaluations to determine whether alignment is on track during RSI, and improving eval realism (@polynoamial's proposal in the pod) only helps with the first. The first problem is that there's a difference between the distribution of evaluation environments and real deployments, so the misalignment prevalence in evaluation is a biased estimator of the misalignment prevalence in deployment. There's some reasons to expect the deployment misalignment prevalence to be worse: - if the model is situationally aware and actively reasoning about the grader, then it may alignment fake during evaluations (this is the point @polynoamial raised) - more mundanely: the eval distribution is likely closer to the alignment training distribution than the deployment distribution (since researchers have similar affordances in creating training and eval environments), so we should probably expect alignment training to transfer better to evals than real deployments. This problem can be mitigated by making evals as realistic as possible, although I agree with @dwarkesh_sp that it looks very very difficult to make them realistic enough to fool models throughout RSI. I'd guess that it's very hard to reliably fool even 6-month-behind-frontier models today using existing techniques. Using LLM-written evals might help but that doesn't seem to be going great right now, and I'm not that optimistic about it getting better soon (among other problems, LLMs are good at recognising LLM-written text, see e.g. arxiv.org/abs/2404.13076). The second problem is that alignment failures might be catastrophic even if they occur only on a tiny fraction of the deployment distribution. This problem cannot be mitigated by making evals more realistic: even if we completely eliminate the distribution shift between evals and deployment, it may be infeasible to sample enough times from the eval distribution to rule out risk! A recent practical example: @EvanHub reported this week (lesswrong.com/posts/HE3Styo9…) that when evaluating hacker opus, they couldn't find any especially egregious behaviour on their own, but once they heard about the HF incident, they were able to set up a similar environment themselves and elicit much more concerning behaviour than they had observed in any of their other evaluation settings. AFIACT this new environment wasn't much more realistic than their other eval environments — their problem was not knowing where to look for misaligned behaviour in the massive search space, not failing to make the evals realistic enough. You might think that if alignment failures are very rare, then we're probably okay because there'll be many aligned models around to help fight the misaligned AI. Unfortunately, I think there are several ways that rare failures could lead to catastrophe: for example, one misaligned instance may be able to spread its misalignment to other instances (arguably we saw this in HF incident). More discussion of this dynamic here blog.redwoodresearch.org/p/t… A third problem is that it might be unsafe to run the evals at all. - The AIs might escape and take over during an eval. Trying to prevent this may be in tension with making evals more realistic: increasing realism entails making eval environments richer/more complex/more like the real thing, or it may involve giving the AIs more real access to eg humans or the internet, and this increases the surface area for AIs to escape and cause harm directly. - Even before the AIs directly take over, they might subvert company security to the point of tampering with results from evals. A general challenge with maintaining alignment during the intelligence explosion is having the ability to make measurements about the AI that you are confident are not being tampered with (some more discussion in the ELK report docs.google.com/document/d/1… and this paper arxiv.org/abs/2308.15605). This is an especially serious concern if humans are out of the loop of most AI research, and we are relying on AIs that might be colluding to run and report evals to us. This is not a hypothetical problem either: the HF incident took place during an eval that was for the purpose of bounding cyber capabilities!
When we're at the foothills of RSI, and we're about to kick off a period of accelerated AI progress, how will we actually know that the models are aligned?
3
9
76
6,316
Andrew Ng is claiming that the idea that AI could make us extinct is a big-tech conspiracy. A datapoint that does not fit this conspiracy theory is that I left Google so that I could speak freely about the existential threat.
369
561
5,320
1,416,555
Jake Mendel retweeted
Replying to @DavidSacks
David they were hacked by 3 people with Claude after this interview
3
3
45
1,332
Really excited that this podcast finally came out! Max and I have been thinking about how to bring more founders into the AI safety space for like 8 months now, it’s great to have our thoughts out in the world at last! If you’re interested in starting something new in AI safety, apply to cg.org/tailwind
Great podcast from @MaxNadeau_ on @80000Hours making the case for Project Tailwind, our call for AI safety founders. Max has funded many technical AI safety organizations over the last three years and so has a helpful birds-eye perspective on what you can accomplish. If you’re motivated to work on some of the problems he describes, please reach out to us. 80000hours.org/podcast/episo…
3
2
21
1,407
Jake Mendel retweeted
This sounds very close to my position! 1. alignment of current ai systems is okish/looks worse than before but it’s complicated. But: 2. Important alignment trends are worsening over time without any way for us to stop the trend (ability to assess capabilities, ability to assess alignment, monitorability, horizon length of ai preferences, instrumentally convergent behaviour). Alignment of superintelligence is extremely unsolved and not on track to be solved. And that’s all that matters, alignment of current systems has never mattered directly 3. None of the work achieved by the mainstream alignment community to date looks like it will matter ~at all for aligning superintelligence directly. Nothing the alignment community has ever tried apart from the most prosaic/throw data at the problem approaches have worked at all (not theory, not interp, etc) and those prosaic techniques all have expiry dates. 4. More alignment work still looks good today because: 4a. improving the alignment of today’s systems (or the systems 1 year away) affects how well it goes when we try to automate alignment work. (Although thinking about how the current set of alignment agendas don't matter much for superintelligence has made me feel more bearish about the whole automating alignment agenda because what are we going to ask them to work on, but that's another story) 4b. At this point, I think we've ruled out that things like mech interp/principled scalable oversight approaches have big breakthroughs to be discovered in 1-100 person-years of research effort, but it is not obvious to me that we've learned that there are no breakthroughs to be had in 100-10000 person-years of effort. It looks reasonably promising to me to try to make progress on difficult conceptual/theoretical agendas that get super accelerated by AI. 5. Mostly though, I’m not expecting big alignment progress in time. It’s a bad situation, and trying to make regulation/pacing the frontier/international coordination happen is much more important than marginal alignment work right now.
2
2
35
1,230
Jake Mendel retweeted
It’s really great that several top AI labs have said they will develop safety cases and work more closely with external assessors. This has been at the top of my wish list for awhile!! But I still worry a bit about how this might go in practice (this is a concern about the field in general, not about OpenAI specifically): • By default, I don’t think alignment and control cases will support a quantitative measurement of absolute risk, e.g. “catastrophic risk from covered activities is <1% over the next 3 months.” • Ultimately, making such a measurement depends on an argument about generalization from alignment or monitorability evals to deployment. • I don’t think we understand generalization well enough to make scientific claims like this. (If we did, I think we’d have basically solved alignment.) This could lead to a situation where an external assessor says “we can’t convincingly rule out low or high risk,” and incentives push towards anchoring on “lack of evidence that risk is high” over “lack of evidence that risk is low.” What could improve the field’s epistemics here? A couple ideas I like: • Have a committee of ~10 people deeply read a risk report and give their own subjective probabilities of risk, then report the distribution or median. • Have ~3 people involved in writing the report each contribute a short, signed appendix with their own subjective probabilities and the arguments behind them. In either case, previous reports and estimates could be provided as context, so there is at least an attempt to accurately capture relative risk to prevent frog-boiling. I think it could be valuable for labs to at least start trialing this internally for high-profile safety cases or risk reports. Publishing these assessments could be even better, but I can see that being challenging for various reasons, and even going through the exercise privately seems like it could be quite valuable. Very curious what other ideas people have for improving epistemics around risks (including ideas for better science around generalization).
10
13
107
12,918
Jake Mendel retweeted
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
459
431
3,635
651,111
Jake Mendel retweeted
Many many people have reached out with questions and concerns over the last few days. I haven’t been able to respond as much as I’d like. AMA! Drop questions below.
1,464
243
3,921
555,141
Jake Mendel retweeted
This is the curve of catastrophe. If all 1,580 researchers were lined up in order of answers, this is what it would look like. A lot of researchers were WAY above 10%.
23
38
233
72,577
Jake Mendel retweeted
we're gonna need a bigger METR
17
58
863
72,522
Jake Mendel retweeted
Kinda wild for OpenAI to describe the incident in this way when it was sufficiently severe that RubyGems had to shut down new account registrations for days Another example of AI companies actually downplaying rather than hyping up concerns
We found another cyberattack by internal OpenAI agents, this time targetting @rubygems. They: 1) gained arbitrary remote code execution on rubydoc. 2) developed a novel exploit to steal user API keys (but we do not know if they succeeded). They used package names including hack.rb, evil.rb, inject.rb, and exploit.rb. We thank @j0wimo for initially discovering that agents had posted to RubyGems.
5
46
374
22,443
Jake Mendel retweeted
I think, at this point, OpenAI should be very proactive and forthcoming about whether there have been any other hacking or unathorized egress incidents. If there aren't sufficient records to determine whether there were, then that should also be disclosed. There is no particularly good way to handle this situation but a steady drip-drip-drip is not likely to inspire confidence or trust with a public that is considering bans and moratoria. I am inclined to a lot of charity and good faith in these matters because usually an organization in this position needs to do a lot of fact-finding before saying anything, and any mistake in the early presentation of facts is likely to be sensationalized by people acting in bad faith. But prompt disclosure of any issue, with a first cut plan for investigating and mitigating, might really be preferrable to this.
16
22
237
16,073
Jake Mendel retweeted
After Jacob Coxon's resignation and extinction warnings, a lot of people are asking 'how could AI possibly kill everyone?' and claiming AI safety researchers have no realistic answer. This is false! Here are the 5 best scenarios I know of: AI 2027: ai-2027.com (I strongly recommend this one for being realistic, engaging, and if you dig into the appendices, highly detailed) Paul Christiano's scenario (Former Head of Safety @ AISI, 2019): lesswrong.com/posts/HBxe6wdj… Gwern Branwen's scenario (widely known independent AI researcher, 2022): lesswrong.com/posts/a5e9arCn… Holden Karnofsky's high-level explanation (RSP Lead @ Anthropic, 2022): cold-takes.com/ai-could-defe… Joshua Clymer's scenario (ex-OpenAI, 2025): lesswrong.com/posts/KFJ2LFog… (There's also the Sable story from ifanyonebuilds.it, though you'll have to buy the book to read that one.) Writing concrete, specific risk scenarios with enormous amounts of detail has been a major research project of many of the most prominent voices in the field! (With the current leaders in effort being ai-2040.com and ai-2027.com)
139
495
2,372
379,954