OpenAI researcher
@boazbaraktcs says the Hugging Face incident didn’t change his view on alignment, and points to the long-term trend that worries him more:
"Every incident or every bump between one version or the next, you tend to overweight it."
"If you fix any alignment eval, then like we do for all evals, we are getting better at it and we'll quickly saturate it. But since model capabilities are growing, that's not good enough."
"I still think fundamentally that we are improving in alignment, but nowhere near as fast enough as the level of capabilities grows. Which is why I think pacing is a good idea, because we do need alignment and safety to catch up."
@OpenAI