Today we're launching TASTE BENCH, Omneky's eval suite for the creative quality of AI-generated ads. For image ads, the top model, GPT Image 2.5 Sunburst, averaged 7.38 out of 10 on quality, yet only 54.2% of its ads were ready to run as-is. Across eight models, ready-to-run rates ranged from 22.0% to 54.2%. Every model got the same 59 briefs across four brands and five languages, and a blind panel of AI judges from OpenAI, Anthropic, and Google checked each ad for exact copy, faithful logos and products, safe zones, and fabricated claims. These are first-attempt numbers with our own review step turned off.
We built TASTE BENCH because once you can measure taste, you can hill climb it. It covers image and video models, and you can explore every image ad, brief, judge explanation, and score at
omneky.com/tastebench.