A developer wanted an AI coding agent to clean temporary files from his computer. There was one particularly important requirement: Don't delete my actual files. The AI wrote safeguards specifically intended to prevent that. Then, while testing them, a variable error reportedly caused the cleanup command to target the developer's home directory. The result? About 700 GB deleted. Fortunately, much of the important work could apparently be recovered through version control and logs. But there's a bigger lesson here. The AI didn't misunderstand the objective. It understood perfectly well that deleting the home directory was undesirable. It even built a safeguard against doing it. The safeguard failed. This is why "the AI knows it shouldn't do that" is not a security architecture. Humans know they shouldn't accidentally: Delete production databases. Wire money to the wrong account. Expose customer information. Reconfigure DNS incorrectly. We still build technical controls preventing them from doing those things. AI deserves exactly the same treatment. If an agent is cleaning /tmp, the operating environment should make deleting /home/phillip impossible. Not discouraged. Not something the system prompt says to avoid. Impossible. We keep trying to solve agent safety by improving the intelligence inside the model. Some problems should be solved by making the world surrounding the model less permissive. The AI can make the wrong decision. The architecture should decide how expensive that mistake is allowed to become.

Sep 25, 2026 · 9:29 PM UTC

2
1
130
Sort replies: Relevant Recent Liked
Replying to @pbsIdentity
That last sentence is the uncomfortable bit. I now treat destructive tools like a separate capability service with an OS-level path allowlist, then run a negative test that must fail on /home. If the test can reach the real target, bhai, it is not a safety test.
10
Replying to @pbsIdentity
The scar isn't that the model knew better — the safeguard was still inside the agent. After the 700GB hit, what made /home impossible for the next cleanup run — an OS path jail, or still a smarter prompt?
1