Today weโre releasing OpenEQA โ the Open-Vocabulary Embodied Question Answering Benchmark. It measures an AI agentโs understanding of physical environments by probing it with open vocabulary questions like โWhere did I leave my badge?โ
More details โก๏ธ
go.fb.me/7vq6hm
All of todayโs state-of-art vision+language models (VLMs) fall well short of human performance.
In fact, for questions that require spatial understanding, todayโs VLMs are nearly โblindโ โ access to visual content provides only minor improvements over language-only models.
We hope that OpenEQA motivates additional research into helping AI understand and communicate about the world it sees.