We were concerned with the quality of LVBench, a leading benchmark for long video understanding, so we contracted a data company to fix it. It showed us we actually still beat Gemini on LVBench.
This result simply reaffirmed that our hill climbing on solving real-world tasks also transfers to the popular benchmarks.
These results are from last month, so stay tuned for updates. Video evals and real-world tasks are a sneak preview of what's coming soon to use[.]vision 👀...