We spent 9 weeks testing agent reliability during software tool use. Nearly 60% of claims accepted as correct could not be backed by evidence from recorded tool calls. The answer looked right. The evidence wasn't. 🧵
2
5
6
130
We tested 15 models across 180 runs and 570 claims whilst using @scale_ai s MCP Atlas. In 161 claims, the final answer and the agent's recorded tool calls told different stories. 109 claims appeared in the answer without evidence. 52 had evidence that the agent failed to use.
1
55
That is 109 out of 183 claims accepted as correct. Nearly 60%. We cannot say whether the model guessed, remembered the answer or found it elsewhere. We can only say its own recorded work did not support the answer.
1
15
The reverse matters just as much. For 52 claims, the agent had already retrieved the supporting evidence but failed to use or present it correctly. The retrieval work was done. A potentially correct answer was left on the table.
1
7
Hard tasks exposed the gap. Answer scoring passed 7 of 60 runs. Complete trace support passed 0 of 60. The answers looked correct, but the recorded execution could not support every required claim.

Aug 11, 2026 · 2:27 PM UTC

1
5
We believe agents should capture evidence natively as they work. SourceryKit captures tool calls in a state of the art verifiable database, allowing agents to evaluate themselves and downstream agents to independently verify their answers.
1
12
SourceryKit detected 100% of the covered deterministic grounding errors in this study. Next, we will test how much accuracy agents can recover through answer healing and targeted retries. Read the study: provably.ai/blogs/The-Agent-…
9
Sort replies: Relevant Recent Liked