We need to be running these types of experiments to better understand coding agents empirically - evals are not the only method!
Confirms the idea that models are much better at adding than reducing entropy.
I would even question the conclusion that this refactoring was successful (due to reducing token usage) - breaking down into better abstractions is good but I'm deeply suspicious when in doesn't reduce LOCs.
Last edited Jul 30, 2026 · 4:34 PM UTC
3
198




