7 AIs, the same Flutter repo, the same audit prompt.
🏆 DeepSeek V4 Flash took the lead with a 92.9 report quality score and an 88.8 overall score, using 289K tokens in 19 minutes.
DeepSeek V4 Pro finished in just 7 minutes and scored 88.5 overall, only 0.3 points behind the leader.
GPT-5.6 Terra produced the second-best report quality score at 89.6 and ranked third overall with 85.3.
💸 Ox Alpha, currently getting attention for being free, scored 82.8 overall at $0 cost, making it one of the most interesting results in terms of cost efficiency.
⚠️ GPT-5.6 Luna (XHigh) used the fewest tokens in the test at only 117K, but still finished last with a 69.7 overall score.
The task: read all code and Markdown files in the project, make no changes, and produce a detailed report covering bugs, code-documentation inconsistencies, optimization issues, and other technical problems.
✅ flutter analyze: 0 issues
✅ flutter test: 126/126 passed
The reports were evaluated by GPT-5.6 Sol.
Detailed comparison tables are in the next post.