In Saturn’s test of 18 AI models, financial answers were wrong 57% of the time on average. For complex questions, that rose to 88%.
Financial tasks expose a difficult property of AI: an answer can be internally coherent while resting on a false premise.
The arithmetic can be correct while the tax rule is outdated. The source can be real while applying to the wrong jurisdiction. Every subsequent step can look reasonable because it inherits the same initial error.
Give that output access to tools, accounts and payment rails, and the error can propagate into an action.
This is why verification needs its own place in the architecture.
Before execution, a financial agent should have to pass several distinct checks:
• Evidence: do the sources actually support the claims?
• Context: are those sources current and applicable to this user?
• Computation: can the numbers be reproduced with deterministic tools?
• Authorization: is the proposed action within the user’s permissions and limits?
Asking the same model to reconsider its answer provides limited independence. Agreement between multiple models is also insufficient when they share the same blind spots.
Verification needs external evidence, explicit constraints and a way to stop execution when checks fail.
The architectural question for agentic finance is becoming unavoidable:
What has to be independently checked before an output is allowed to move money?