Two teams can build on the exact same model, and one burns 40x more tokens to get the same work done.
Same model. So what is actually different?
Most of the conversation right now is about which model is best. It is the wrong argument.
Look at what a model upgrade actually gets you. Somewhere between 4% and 10% better results, at about 1.3x the token cost. Real, but small, and you keep paying that 1.3x every month.
Now try the other lever. Leave the model alone and fix the layer around it.
That layer is the harness. Everything the model does not do on its own. What actions it can take. What information it sees, and in what order.
Get that right and the same model runs on up to 40x fewer tokens, with up to 10x less time spent going in circles.
Two costs hide in a badly built harness.
1⃣ The first is context. A sloppy harness resends everything on every call, so you pay for tokens the model did not need to see.
2⃣ The second is idle turns. The agent takes turns where it does nothing useful, and each one can still trigger another inference. Tokens and latency, no progress.
Coding got there first. Claude Code and Codex put real engineering into the harness layer: the agent loop, context management, tool execution, sandboxes. The harness understands code and the specific ways things break.
Science is earlier on that curve. There are research agents, but domain harnesses for science are far less mature than coding harnesses, and it shows in the reliability.
A research harness is a harder build. It has to hold together 220M+ scholarly sources, 70+ connected databases, and 393 research workflows, and still hand back work a researcher can defend.
That is what we are building at SciSpace.