What OpenAI built here is incredibly Machine Learning. Define a large set of metrics a class of experts need to agree on, cross reference against user feedback, build scoring. 🙃 The nuances are legion.
But generally, I might be a good example of how this could lead to optimizing for wrong metrics if taken superficially. Hear me out:
I really don't like most therapeutic approaches. The first successful uninterrupted 2 year therapy I had in my life was from 29..31 years old. Why? Because in order for me to accept being tinkered upon mentally, it took me understanding the mechanics of my soul, not for someone to ask me questions and guide me. That just triggers my analysis of their analysis.
I'd react negatively to the best-on-average approach indicated to be taken by the headlines results, and I think I'm thankful for clinicians' and users' ratings diverging here.
Because that reflects the reality of "one size fits all" approaches to problems.
The benchmark here is necessary and good work, but studying how it is built and how grading works in it will take me a while, with the likelihood of me emerging with a large set of criticisms of how to apply its insights feeling quite high.
We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench.
This new open benchmark was built with input from more than 80 mental health clinicians.
We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work.
openai.com/index/introducing…