Your one stop shop for real human data annotations, feedback and opinions.

AI Week is coming up in Zurich; Rapidata is taking care of the afterparty ;). If you're working on multimodal evals/post-training and want to mingle with fellow researchers, you probably don't want to miss our Apero.
1
3
48
Breaking: @AnthropicAI's opus 5.5 shoots up to 3rd place in SVG generation. Making it by far the best generalist model. @OpenAI's gpt 6 luna and sol also make a debut but fall short of the frontier.
4
323
Visual realism ≠ physics. Based on Physics-IQ, a @GoogleDeepMind paper, we benchmarked 26 world models on 65 real-world physics scenarios, 450K+ human judgements Gemini Omni 1.1 Flash takes #1, followed by Minimax H3. Outputs & votes are visible on benchmark.ai
2
4
11
543
Here an example, where you can see the old Omni Flash model fail and the latest one succeed.
1
24
Released a few hours ago, @QuiverAI’s 2 latest models just entered our SVG benchmark with 2.5M+ human votes. Last week, Astra had a considerable ELO gap over the field. This week, Quiver did the same, ranking number 1 across all dimensions. 👀 Check it out on benchmark.ai/svg
2
9
2,368
Complex camera trajectories expose the gap between video models and world models. Sana WM from @nvidia + other world models move up. Typical video models start lagging. Check our camera motion benchmark: 14 models, 600 prompts, 145K human responses. Fully transparent and reproducible. benchmark.ai/camera-movement…
2
4
11
2,673
When easy camera trajectories are included, the video models move up. You can select the trajectory steps as filter on the Benchmark.ai menu. benchmark.ai/camera-movement
2
114
Sota-ing the Sota that you just Sota-ed already today @OpenAI. Sunburst, by far number 1, according to almost 5M human judgments. benchmark.ai/image
1
2
155
Released a few hours ago, we already ran GPT Image 2.5 (Flare) from @OpenAI through our new T2I benchmark, including 4.8M human judgments over 1.5K prompts. #1 among public models, again. Congrats 🫡 benchmark.ai/image
1
3
155
No words for Astra. Congrats @OpenAI team 🫡 .We rarely see a gap this large. @GoogleDeepMind’s Gemini 3.8 Flash also doesn't go unnoticed, jumping to #3 on Alignment.
1
3
96
Arena style leaderboards are the industry norm for visual AI. What is being measured changes as certain prompts are trendy. Say your model is amazing at UI elements. Just make that go viral while your model is being tested in arena.ai, wham bam you're Nr1.
1
86
Millions are spent training models for specific capabilities, then judging them with: “Which output do you prefer?” 💸 We need more granularity + transparency. So here comes benchmark.ai. First benchmark: SVG generation. 42 models, 1.9M+ human judgements. Prompts, outputs, match-ups & methodology are public. What benchmark do you want to see next?
1
3
13
27,403
Human feedback can arrive fast enough, without sacrificing quality, to become "The reward signal" for model training. If I tell you this, you'd probably wonder about the quality of that human feedback. So let's dive in 😉 AI labs currently have two main options: approximate human judgements with a reward model, or collect human feedback asynchronously through traditional crowdsourcing platforms. Let's compare Rapidata through typical crowdsourcing. We sent the same 300 visual tasks to @Rapidata and Prolific and published all 18,000+ individual responses. The result is the following: Prolific participants were more accurate individually: 98.0% vs. 91.6%. But once responses were aggregated, both crowds reached essentially perfect final-label accuracy. On this experiment, Rapidata delivered those labels: → 70× faster (15K answers in 2 minutes vs 3K in 66 minutes) → 8× cheaper → Both with >99.9% accuracy with enough answers per item At this speed, human feedback does not have to remain a slow and asynchronous annotation step, nor has to be approximated. It can be added as a direct reward signal during post-training or, for some training loops, replace the reward model entirely. We first explored this application while supporting the post-training of one of the leading image models out there. Full methodology, limitations, results, and raw dataset in comments.
5
1
11
264
One of the few models where the generation took less time than our evaluation! 🔥
Today, we're introducing P-Image-Ideogram: a family of Pareto-optimal image models with the best quality-speed-cost trade-off, co-developed with @PrunaAI. Four quality modes. Native 1K and 2K generation. From $0.003 per image. Live now on the API and all our partner platforms. Give it a try: ideogram.ai/tools/p-image-id…
3
1
11
868
Rapidata retweeted
Replying to @ideogram_ai
Human votes with @RapidataAI places P-Image-Ideogram as optimal on quality-speed and quality-price.
4
12
193