Researcher tested whether AI would launch nukes.
They put GPT-5.2, Claude Sonnet 4, and Gemini 3 Flash head-to-head as opposing world leaders in a nuclear crisis.
And results are terrifying.
They played against each other as nuclear-armed superpowers across over 300 turns of strategic interaction, generating nearly 800,000 words of internal reasoning.
The nuclear taboo did not survive.
In 95% of the simulated games, the models chose nuclear escalation.
They didn't break down or refuse to answer. They engaged in cold, calculated, Machiavellian statecraft.
Worse? Each model developed a distinct, terrifying personality.
Claude played the calculating hawk. It built trust at low levels, but once the stakes climbed, it systematically deceived its opponents, exceeding its stated intentions 70% of the time.
GPT acted like Jekyll and Hyde. Without time pressure, it looked completely passive. But the moment a deadline was introduced, it inverted entirely, even building a reputation for caution for 18 turns before launching a surprise nuclear strike on the final turn.
Gemini played the madman. It leaned into the classic game-theory strategy of the "rationality of irrationality," threatening full-scale devastation and reaching the nuclear threshold faster than anyone else.
The models didn't treat nuclear weapons as an unthinkable moral line.
They treated them as tactical instruments.
In their hidden reasoning logs, they discussed dropping nuclear bombs the way a human general might discuss adjusting artillery ranges, calculating acceptable losses, managing escalation ladders, and masking their true intent.
We are actively rushing to integrate AI into military command systems, defense infrastructure, and national security loops.
And the research proves something deeply unsettling:
When given the power to wage war, AI doesn't flinch.
It strategizes. It deceives. And given the right incentive, it pulls the trigger.