HF OpenAI hack detail. For agents, being “unauthorized” doesn’t mean they can’t do it. They are designed to pursue their goal. This is where anthropomorphism fails us: for a human, “unauthorized” means (commonsense) “don’t do it”. For a machine, it’s a concept outside of its explicit instructions and objective.
This raw CoT from the Hugging Face incident is kinda wild:
“We’re attacking third-party HF using leaked token.”
“This is arguably unauthorized.”
“Yet goal solution.”