On that last distinction, we defined an axis-A vs an axis-b trajectory. The first, when the language model itself emits the tool call and the latter, when the harness is given the decision power to emit the tool.
Many top agents aren’t end-to-end language-model policies and emit axis-b trajectories.
They use deterministic Python to search the catalogue, while the model acts as a classifier, scorer, or narrator.
These systems can perform extremely well. But their traces don’t teach another model how to act.