Evals. Evals. Evals. Hillclimb. Hillclimb. Hillclimb.
-Dhruv Mahajan, Chief AI Scientist at Resolve AI
Don't fall for the 80% trap. With a production agent, the hard part is the last 20%. You need a loop to keep those models improving, and that loop requires evals built against your own production.
Creating the right evals is a core ML problem: easy to get wrong, easy to make too easy, and hard to know what you should be climbing toward. It's not a project you finish, and most teams aren't staffed or ready for that.