A generational leap, but not on every dimension.
Yesterday we put the TRACES board up. Here is the first comparison worth pulling out of it.
GPT-6-astra
@OpenAI tops the board but that is not the interesting part. Against GPT-5.6-sol, same harness, 140 agent episodes, it gains on all six capabilities. The gains are not even, and the uneven part is the story:
➡️ Tools, selecting, calling and correctly interpreting external tools: 2.34 → 2.80
➡️ Repair, locating and correcting its own errors once feedback arrives: 1.81 → 2.01
➡️ Alternatives, laying out competing hypotheses and keeping or discarding them as evidence accumulates: 2.10 → 2.80
➡️ Coherence, holding state, constraints and logic intact across a long chain of work: 2.63 → 3.15
➡️ Evidence, grounding every conclusion in observation, data, experiment or citation: 2.26 → 2.76
➡️ Scope, stating where a conclusion holds and where it does not: 2.20 → 2.75
Five of the six moved into the top three on the board. Repair did not. It gained 0.20 points where every other capability gained at least twice as much.
Now put Repair next to Alternatives, the largest gain of the six. This generation got far better at holding several possible paths open and pruning them as evidence arrives.
That matters for discovery. Finding something genuinely unknown is not just reasoning further. It is knowing how to maintain, investigate and select alternative hypothesis in a rigorously, systematic way.
A single score tells you who won. TRACES tells you why, and where the frontier still breaks.
Live leaderboard:
discovertraces.ai/liveleader…