Infinibench-v0: part of the benchmark was created by a continuously exploring agent, rewarded for identifying new failure modes in VLMs.
It produces non-trivial, out-of-distribution problems — trivially easy for humans, yet SOTA VLMs fail — and diagnoses failure modes that were previously unknown.