Human feedback can arrive fast enough, without sacrificing quality, to become "The reward signal" for model training.
If I tell you this, you'd probably wonder about the quality of that human feedback. So let's dive in 😉
AI labs currently have two main options: approximate human judgements with a reward model, or collect human feedback asynchronously through traditional crowdsourcing platforms. Let's compare Rapidata through typical crowdsourcing.
We sent the same 300 visual tasks to
@Rapidata and Prolific and published all 18,000+ individual responses.
The result is the following:
Prolific participants were more accurate individually: 98.0% vs. 91.6%.
But once responses were aggregated, both crowds reached essentially perfect final-label accuracy.
On this experiment, Rapidata delivered those labels:
→ 70× faster (15K answers in 2 minutes vs 3K in 66 minutes)
→ 8× cheaper
→ Both with >99.9% accuracy with enough answers per item
At this speed, human feedback does not have to remain a slow and asynchronous annotation step, nor has to be approximated.
It can be added as a direct reward signal during post-training or, for some training loops, replace the reward model entirely.
We first explored this application while supporting the post-training of one of the leading image models out there.
Full methodology, limitations, results, and raw dataset in comments.