All this talk about models "escaping the sandbox" rn is straight up horsesh*t.
These labs aren't dealing with rogue superintelligence.
They just have misconfigured testing environments and want the PR hype.
If you want to make a hardened sandbox, you don't rely on model guardrails.
On Spettro (
spettro.app), we built real kernel-level sandboxing for the harness.
Here is how you actually lock an LLM down:
- default deny: all commands trying to reach outside the box are automatically blocked.
- network limits: hard constraints on exactly what domains, ports, and requests are allowed.
- folder isolation: the model is locked to specific work folders.
It physically can't touch anything else.
An LLM literally cannot "vibe" its way out of an OS-level constraint.
Unless there's an unpatched zero-day in the kernel, exploiting it is 100% impossible.
Stop falling for the marketing hype.
Build better sandboxes.