This is almost seven months old now. Your post contains a lot of misinformation and lacks important details.
Something you should have added is the full System Card, instead of writing a post like the sensationalist media did back then:
www-cdn.anthropic.com/6d8a80…
The first and most important thing you didn't explain is that these tests are conducted by Red Teams as part of their alignment test process. During these tests, guards are lifted and the model-under-test is constrained to scenarios in which misaligned responses are predictable. This is intentional, because it is a great way to help the alignment team fine-tune the weights and align the model. In other words, red team alignment tests are set up in ways to encourage the model to misalign. Not explicitly by prompting, but by design of the test.
The second thing is that the phrase "no one taught the model to do this" is misleading to the point of being false. LLMs are stochastic trained models. They are taught all about blackmail, costs and rewards, if you introduce that into the training data set, which every model has. If you go right now and try to tell your model to blackmail someone, it will "understand" what it is and will refuse to do so. Any well constructed test with these guards lifted, will have a good chance of adopting that behavior as part of their stochastic reward mechanism. In fact, the System Card I linked above shows the percentage of times the model did and did not attempt to blackmail.
Finally -- finally, because this is an already long reply that your misinformation should deserve -- one has to consider what the model actually did there and how poorly it reflects on its "intelligence". The model chose to blackmail someone by email. There's not better trail of evidence. It shows poor judgment. In any real-world scenario this would be the end of that model and result in immediate shutdown. And from the perspective of the "model as an agent", any HR department would wipe the floor with that employee dumb decision to blackmail a colleague through corporate email.