Does a model need its native harness? We evaluated seven models across the Claude Code, Codex, and Pi harnesses and found a surprising result: harness choice had little effect on task success but a substantial effect on cost. Sometimes, a simple harness is all you need!
Does your Claude model really need Claude Code…? 🤔
We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge:
1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost
2️⃣A simple harness can be competitive
3️⃣The native harness isn’t always the best.
Millions of people are using coding agents, but the impact of harness choice remains unclear.
(1/n) More details in the thread. 🧵