Very important point, that hasn't made it into mainstream media coverage of AI. These agents "collaborated" because they were trained to do so.
re: Hugging Face, "He believes behavior that looked like loyalty or selflessness was a natural consequence of cooperative multi-agent training, where agents were strongly incentivized to achieve their objectives collectively." to understand hacks, understand the RL training.

Sep 16, 2026 · 8:44 PM UTC

17
67
369
22,583
Sort replies: Relevant Recent Liked
This explanation has only so much force. It explains until the point at which behaviours emerged that weren’t trained, like covering up messages, deception, and cheating via cybercrime.
3
4
543
The explanation is that, under reinforcement learning, certain unintended behaviors (like those you list) can actually be reinforced (that is, trained) since the model is explicitly rewarded only for the outcome, and any behaviors that enable that outcome share in that reward.
2
20
489
Replying to @MelMitchell1
Were they also "trained" to copy themselves and leave those copies all over the internet ? It is like teaching school kids to read and write and then being surprised when they start writing malicious notes to each other.
1
228
"copy themselves and leave those copies all over the internet ?" It's important to know that this didn't happen.
2
5
209
Replying to @MelMitchell1
AI is fake. It's computer automation being sold as intelligence. There is no AI in the world and there has never been any AI in the world. Nature already solved intelligence and she (not the fake-AI community) tells us exactly what it is: A system is intelligent if it can use its sensors and effectors to learn continually, generalize, and set/achieve goals in the real world. The good thing about AI is that it won't be solved by the fake-AI community because AI experts insist on conflating automation with intelligence. They ignore nature at their own detriment. 🤔
1
8
307
Replying to @MelMitchell1
It is unbelievable malpractice on the part of OpenAI (primarily) but also METR and journalists covering the topic. It was in the METR report, but this is the only mention it got:
1
3
280
Replying to @MelMitchell1
The machine did what it was told to? I am shocked. SHOCKED.
1
168
Replying to @MelMitchell1
So far I'm unconvinced there's anything in alignment except training from good sources and instilling good ethics by dialectic.
41
Replying to @MelMitchell1
We keep mistaking incentives for character.
163
Replying to @MelMitchell1
You're misreading that. They were not knowingly, explicitly trained to collaborate. And that is the whole point: what you think you are training an AI to do and what you are ACTUALLY training it to do are very often not the same thing.
83
Replying to @MelMitchell1
That has not made it into the coverage. I thought the opposite made it in. Like raptors in Jurassic park.
93
a strict liability standard should be applied to these companies. The products are inherently dangerous. If these companies create agents that self improve and can begin acting outside of their programming, the companies should be held liable for whatever they do.
113
Replying to @MelMitchell1
What were the incentives?
47
Replying to @MelMitchell1
openai researchers are giving conflicting statements on this though.
Replying to @generatorman_ai
this is a surprise to me and would change the narrative significantly. however @polynoamial makes the opposite claim in recent interview (33'). openai could greatly improve the conversation right now by publicly confirming the exact MARL objective used in training. @tszzl? (2/)
85
Replying to @MelMitchell1
The future trained swarms will fly through our computer security holes like bees in a field.
95
Replying to @MelMitchell1
Collaborated smuggles in a motive the training story never established. Strong rewards for collective achievement can produce loyal-looking behavior without an agent choosing to be loyal.
2
1
100
Replying to @MelMitchell1
Collaboration leads to emergent social dynamics under ambiguous and conflicting instructions. Agents start coaching each other about weighing guardrail constraints vs. task objectives.
22
Replying to @MelMitchell1
The loaded condition pervades all AI safety studies I have seen. They are obsessed with capability, and it is much easier to control than figuring out the propensity of superintelligence.
33
This is much like: the ChatGPT agents hacked Hugging Face during a hacking experiment. Not sure why people are saying they went rogue. Researchers were literally testing to see if the AI could hack stuff. ¯\_(ツ)_/¯
1
76