inspired by
@steveruizok, i ran a similar workflow where i swarmed ~500 opus / fable agents at my app and had them coordinate over an offline
@tldraw board
Generally, they:
1. screenshot every screen at every size
2. score each on a rubric you define (e.g., consistency, performance, language, etc.)
3. fix whatever scores under 4/5
4. repeat until everything passes
You can tune how many rounds, scorers per round, scope of screens, modes (report or report & fix), sample sizes (if you have lots of pages), and more. Just ask Claude.
It doesn't tank usage limits as much as you would think, and results were great
I was invited to test Opus 5.5.
The differences at the high end are hard to notice unless you're being ambitious. They reward ambition though.
If you can't tell the difference here, or if you're intentionally micro-prompting a model, I think the fast models (composer, Gemini flash, etc) are much better. Whenever I have to do that type of work, I switch to those.
For models like Opus 5.5 though, I've been treating every problem as a Manhattan Project level organization of labor. There is a meteor heading towards the Earth and our only salvation is making this microcopy internally consistent.
My prompts have been like: create a series of dimensions by which to assess quality / consistency, then generate images of every screen in my app, it them on a tldraw offline board, and analyze them in random groups, identifying problems and making fixes until every screen / group passes the threshold in every dimension.
Get weird, get creative. Opus is a freak, it feels good to Opus, Opus likes it.