A ton of AI and Robotics developments this week.
Big announcements from Microsoft, Anthropic, Wayve, Cover, OpenAI, Meta, Google DeepMind, Runway, Hedra, ElevenLabs, and Unitree.
Here's everything that happened and how to make sense out of it:
Is your Vision-Language Model really helpful at all times?
Can we instruct them to interact with users during conversations to avoid hallucinations or biased responses?
🍰Take some bites of PIE and MACAROON! We present a benchmark to evaluate LVLMs’ proactive engagement capabilities and train LVLMs to be your real engaged partners through self-iMaginAtion for ContrAstive pReference OptimizatiON.
📃Paper: arxiv.org/abs/2406.14137
On my way to London after an unforgettable week at #CVPR2024. It was my first in-person CVPR and I loved every moment of it. A big thank you to everyone who came to my talk/poster!
Something that puzzles me about the conformal prediction literature: It is very focused on getting O(1/n) coverage bounds -in expectation- over the calibration set, rather than O(1/sqrt{n}) coverage bounds -conditional on- it. But in what applications is this the right guarantee?
This paper seems very interesting: say you train an LLM to play chess using only transcripts of games of players up to 1000 elo. Is it possible that the model plays better than 1000 elo? (i.e. "transcends" the training data performance?). It seems you get something from nothing, and some information theory arguments that this should be impossible were discussed in conversations I had in the past. But this paper shows this can happen: training on 1000 elo game transcripts and getting an LLM that plays at 1500! Further the authors connect to a clean theoretical framework for why: it's ensembling weak learners, where you get "something from nothing" by averaging the independent mistakes of multiple models. The paper argued that you need enough data diversity and careful temperature sampling for the transcendence to occur. I had been thinking along the same lines but didn't think of using chess as a clean measurable way to scientifically measure this. Fantastic work that I'll read I'll more depth.
LLM agents have demonstrated promise in their ability to automate computer tasks, but face challenges with multi-step reasoning and planning. Towards addressing this, we propose an inference-time tree search algorithm for LLM agents to explicitly perform exploration and multi-step planning in interactive web environments.
It is the first tree search algorithm for LLM agents that shows effectiveness on realistic and complex web environments: on the challenging VisualWebArena benchmark, applying our search algorithm on top of a GPT-4o agent yields a 39.7% relative increase in success rate compared to the same baseline without search, setting a state-of-the-art success rate of 26.4%. On WebArena, search also yields a 28.0% relative improvement over a baseline agent, setting a competitive success rate of 19.2%.
I am hiring a Research Scientist for the Google Bard Extensions team in the Bay Area (MTV office).
If you are interested in building the next generation of AI agents, please DM or reach out at goelrahul@google.com.
Happy to share our #ICCV2023 paper on 3D reconstruction from a single image!
In Zero-1-to-3, we teach diffusion models to control the camera viewpoint, which enables novel view synthesis applications.
Website: zero123.cs.columbia.edu
Paper: arxiv.org/abs/2303.11328
🧵(1/n)
Very interesting paper: using generative AI to produce text or images emits 3 to 4 orders of magnitude *less* CO2 than doing it manually or with the help of a computer.
arxiv.org/abs/2303.06219
I ran some scripts on the package index on Hackage and found some interesting stats.
I was going to make a blog post but am not so interested in the backlash so I'll just post them here instead.
C++23: Ranges Improvements and std::generator
C++20 does not provide concrete coroutines, but C++20 provides a framework for implementing coroutines. This changes with C++23. std::generator is the first concrete coroutine.
modernescpp.com/index.php/c2…#cpp#cplusplus#cpp23
Why Can a Unique Index Store the Same Value Twice?
If your unique index includes nullable columns, its behavior to enforce uniqueness may not be what you expect it to do.
👇 Learn More
I have said in the past that AI systems will eventually become the repository of human knowledge and that the only realistic way for such a thing to exist is through crowd-sourced training of an open-source base system.
This paper proposes an architecture.
The real promise of of AI is not that it will make us individually more intelligent, but that it will make humanity collectively exponentially more intelligent. Here's how:
homes.cs.washington.edu/~ped…
We got fascinating results in this work!
* we reverse engineer the training set for Copilot/Codex
* we show that data deduplication can sometimes hurt privacy
* we reveal the tokenizer of black-box LLMs
* we reveal other users' test inputs when adv ex defenses are used
When analyzing ML security and privacy you need to study 𝐬𝐲𝐬𝐭𝐞𝐦𝐬, not just models!
Our new paper shows that privacy is way worse when models are deployed in systems that use data cleaners, output filters, etc.
Paper: arxiv.org/abs/2309.05610
Blog: spylab.ai/blog/side-channels…