Do AI models become less honest when they self-report their own misbehavior?
Our Applied Research team studied a key question in chain-of-thought (CoT) monitorability: do models faithfully report their own misbehavior when they are asked to?
We introduce Monitorability Disposition, a meta-level property capturing how willing models are to disclose their misbehavior.
We found that models reported only ~16% of expected misbehavior. Even with incentives to increase disclosure, no model self-reported high-severity violations. When given a choice, models consistently selected the least strict monitor.
In other words: less oversight is the default choice.
💡 Key takeaways:
- We can measure it: Monitorability disposition is a new way to evaluate how transparent models are.
- Not all models behave the same: Even equally capable models vary in how willing they are to admit when they've gone off track.
- It can be strengthened: When a model’s disposition is strong enough, it stays monitorable.