GPT6 is a very significant jump in capabilities, but also an important decrease in monitorability – especially under adversarial evaluation. We give many details about this in the system card.
In my opinion, monitorability and control will likely become a major bottleneck for responsible AI development quite soon, given that risks from a fixed amount of residual misalignment grows together with capabilities. Aligning on acceptable monitorability outcomes and practices which robustly bound misalignment risks is a non-trivial problem, and there may be limited time to resolve uncertainty around this before we need to act.
It should ultimately be in everyone’s interest to agree on shared bounds for monitorability in order to avoid races to the bottom – nobody wants extremely capable models whose alignment properties we don’t understand, and that are reliably able to cause severe real-world harm without being detected.