This is exactly why physical AI needs a different safety standard from software-only systems. A bad output on a screen can often be corrected; a bad action from a robot may cause irreversible physical harm.
The concerning part is that these failures can happen with seemingly simple tasks, not just extreme or adversarial scenarios. Testing models in simulation is useful, but real-world evaluation needs to become a core part of deployment—especially around force, hazardous materials, humans, and unpredictable environments.
The capability is already here. The safety infrastructure needs to catch up.
the hardware embodiment of frontier models like Claude and GPT is the most urgent AI safety problem in front of us today.
we simulated two very simple use cases using claude both in simulation and using robot arms.
in one, claude spilled toxic liquids in a lab.
in another, the force it used to place an animal toy into a basket was strong enough that it could have physically harmed sensitive material—or anything else in its path.
these are simple experiments, using models out of the box today. as researchers increasingly give frontier models arms, legs, and access to the physical world, we need to urgently build and assess guardrails around what these systems can and cannot do.
models escaping sandboxes or compromising enterprise security infrastructure are serious concerns. however, hardware embodiments introduce something fundamentally different: an AI system can make a mistake in the physical world, and the consequences may not be reversible.
this is not a future safety problem. the capabilities exist today.
we’ve released a report detailing this & solutions we propose. link in comments below.