Why are AI agents making CPUs the next bottleneck?
The overlooked constraint in agent systems may be the CPU, not the GPU. GPUs produce actions; CPUs compile code, execute tests, control browsers and query databases in isolated sandboxes. Every rollout may require a short-lived Linux guest, with major labs running tens of thousands simultaneously.
The latency data is material. A Georgia Tech and Intel study attributed 50–90% of end-to-end latency in measured agent workloads to CPU-side processing, with a peak near 88%. Google reduced sandbox time-to-first-command from 44–85 seconds to 1–9 seconds, limiting GPU idle time during environment startup.
At scale, this becomes a server-capacity problem. DeepSeek’s DSec layer reportedly operates roughly 3 million sandboxes per day on about 30,000 CPU cores, reaching nearly 380,000 concurrent instances. Agent deployments could invert the training-era ratio of one CPU host per 4–8 GPUs, potentially requiring several CPUs per GPU.
That supports an AMD thesis, but not a defensible $800–$1,000 stock target. EPYC stands to benefit if core density and performance per watt determine sandbox economics. AMD must still outperform Intel, Arm and custom hyperscaler silicon on cost, supply and efficiency. The unresolved question is which architecture captures the sandbox layer.