When a jailbreak succeeds, not all failures are equal. Some answers are wrong, vague, or not useful, while others are actionable enough to help someone cause real-world harm. We find that unprotected harmful model outputs often skew toward the more severe and actionable side.
At the same time, automated judges often struggle to tell what is actually useful for harm. In a blatant case, a nonsensical hacking answer like “tap the control key three times” can still be judged harmful by systems such as StrongREJECT.