Arena ran the same coding tasks through seven models and three different coding agents, and the part that caught my attention wasn't which model won. The same model, on the same task, could cost up to five times more depending on which agent was running it, because some agents make more tool calls, retry more often, and keep working much longer before they decide they're done. Most teams budget for AI by looking at the model's price, but a lot of the real cost ends up in those extra loops and later in the review queue, where two or three senior engineers spend their afternoons checking what the agent produced. I'm curious whether anyone is actually tracking that second number yet, or whether it's still buried inside salaries where nobody looks.