Agents on Rails: We ran 8 models against 21 atomic tasks to see which were best at writing Rails code. 3 runs each: a bug report, a security finding, a feature request.
The first benchmark report with findings is now live. So: what did we discover?
As of August 2026:
- Most accurate:
@claudeai Opus 5 by
@AnthropicAI (by a hair). Solved 92% of runs (58 of 63). (But for a little more than half the cost, you get almost the same accuracy with
@Kimi_Moonshot.)
- Cheapest:
@OpenAI GPT-5.6 Luna. 73% of runs solved at default medium reasoning effort, and all 63 of its runs cost 90 cents combined.
- Fastest: Luna again, at a median of 3.3 minutes per run task.
- Best combination of all three:
@OpenAI GPT-5.6 Sol. 84% accuracy, costing $0.52 and 5 minutes per run.
Read all the findings in the first full benchmark report from
@evilmartians here:
rubyonrails.org/2026/8/13/ag…