I believe that Anthropic people believe this.
They think only they can build the Aligned Machine God. They believe they are uniquely moral, and only they will be able to Align the Machine God to what is Just and Right.
This is what they actually think. They're not cynically telling scary stories to sell an IPO. They genuinely believe they are the Machine God's chosen people. Ok.
My question is, how on earth do they plan to effectuate this?
- Do they think, once they achieve RSI, other less aligned AI models will simply cease to exist?
- How does "getting there first" imply that Anthropic obtain a permanent global monopoly?
- Do they think open weight models will never be close to the frontier ever again?
- Do they think only one firm can build RSI?
- Do they believe that once Anthropic is good enough, every other AI model will be crowded out and abandoned?
- Do they believe they can align Claude to a single moral attitude, and that will be sufficient for the whole planet?
- Do they believe that model weights can never be stolen or exfiltrated?
- Do they think that every other AI firm on the planet will converge to their attitude on alignment?
- Do they believe that governments worldwide will cooperate to ban all inference and model weights other than GPUs running Anthropic's model, if it becomes necessary?
I don't believe they have good answers to these questions. But it's extremely problematic, because they are speaking as if a local solution (something internal to Anthropic) can solve a global problem (AI x-risk).
Obviously, if sufficiently powerful AI poses a risk to humanity (I don't believe it does), nothing Anthropic does can mitigate this risk. Anthropic can't go around and ransack every other AI lab or persuade people not to use open weights. They can't force every other lab to converge to their approach to safety. If anything, Anthropic's efforts at "alignment" will only push people further towards competing models (this is already the case). Anthropic products have become an overbearing nanny and snitch; no one wants to operate under a regime like that.
Let's say Anthropic does the impossible and actually "aligns" AI to human preferences (this is impossible, obviously, the same reason no one has "solved" law or ethics or even social media moderation; human preferences are heterogenous and mutually exclusive). But let's suppose they did. The researchers did what they said they were going to, they protected humans from AI x-risk, at least when they use Anthropic models.
Now what? Does Anthropic use their superior AI to extinguish OpenAI, Google, Meta, and xAI? Do they go around destroying every open weight model and provider? Do they sabotage every data center serving non-Anthropic compute?
How exactly does Anthropic intend to translate a corporate discovery to a global status quo? Use their AI to demand the US becomes the global superhegemon and forcibly require Anthropic use everywhere? Destroy all GPUs serving open weight models? Or simply persuade every AI lab on the planet to adopt their approach to safety? And persuade people to stop using open weight models?
The problem with the safetyist view that aligning Anthropic's model can in any way "protect" the earth from AI x-risk is that it contains a nested premise: that Anthropic can export their safety criteria to the entire globe and enforce this uniformly, with absolutely no defection.
This is trivially absurd, and this is why the doomers at Anthropic are peddling inanities when they simultaneously claim there is a 10%+ risk of extinction but they are helping alleviate the problem by working at an AI lab.
"But Anthropic is only responsible for what people do with their own model," you might reply.
Reasonable, to be sure. But the researchers aren't saying that. They are saying that it's ok if we are knowingly contributing to a 10%+ chance of global extinction, because we are working to address it. But they're not addressing it. They can only address it locally. They are completely powerless globally.
Jacob has actually done the reasonable thing, which is to leave. But meanwhile you have lunatics like his colleague Evan (and presumably most staff at Anthropic) that maintain with a straight face that they believe their technology might destroy the world, but that they are actually uniquely moral because they are working on Alignment.
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.