Evals are the most important part of building AI systems, but there’s not a lot out there about building them well.
So I hosted a fireside chat on evals with Liam Bush (
@LangChain),
@lotte_verheyden (
@langfuse), Braden Holstege (
@mercor),
@neutralino1 (
@CoreWeave), and
@soumyamohan (Galileo) to walk me through the tricks of the trade - how they actually work in production, where teams get them wrong, and where they're headed over the next 12 months.
We get into:
-online vs offline evals
-why running evals can cost you more than running the actual agent
-what nobody can see after the agent makes a tool call
-why understanding accounting rules might be harder than AGI
-and much, much, more…
Some of my favorite parts and link to full episode in the comments: