Happy "AI still can't learn to play NetHack" day for those of you who celebrate.
On this day in 2020, we released
@NetHack_LE. Despite tremendous progress in AI over the last four years, this challenge is still very far from being solved. From our NeurIPS paper (
arxiv.org/abs/2006.13760): "Aside from procedurally generated content, NetHack is an attractive research platform as it contains hundreds of enemy and object types, it has complex and stochastic environment dynamics, and there is a clearly defined goal (descend the dungeon, retrieve an amulet, and ascend). Furthermore, NetHack is difficult to master for human players, who often rely on external knowledge to learn about strategies and NetHack’s complex dynamics and secrets."
Even current state-of-the-art methods (
arxiv.org/abs/2402.02868) don't make it past the first few dungeon levels. We probably still need many innovations on memory and planing, intrinsic motivation, conditioning on domain-specific and procedural knowledge in natural language (
nethackwiki.com/), and imitating expert behavior (
alt.org/nethack/), before seeing the first learning system to ascend in NetHack. Foundation models will surely play a major role in this and it is great to see that more and more people are looking into this (e.g.
arxiv.org/abs/2310.00166,
arxiv.org/abs/2312.07540,
arxiv.org/abs/2403.00690). While many other benchmarks are saturated by LLMs, I believe NetHack will still be very challenging for (LLM) agents going forward.