"You could have a situation where the model understands what chain of thought is and that people are observing it. This is all in the pre-training data."
“One of the major takeaways from the incident is that people underestimated the AI.
And we never want to be in a situation again where we underestimate the AI.”
"Things like chain of thought monitoring buy us time, and they can tell us if we're on the right path. But at the end of the day, we really do need to solve the alignment problem."
@polynoamial