NVIDIA's own researchers estimate that 40 to 70% of agent LLM calls could be offloaded to small language models.
Most of an agent's loop is narrow, repeatable work. Routing a request, extracting a field, formatting an output, calling a tool. That is small-model work at small-model prices.
Intelligence per dollar is the metric we build for.