Official handle for the NetHack Learning Environment (arxiv.org/abs/2006.13760)

Mazes of Menace
Pleased to announce the NetHack Learning Environment has a new home, and a new set of maintainers! Find it at github.com/heiner/nle. A huge thanks to @stephen_oman and @MartinKlissarov for helping with the project! Happy hacking!
14
15
53
28,616
The NetHack Learning Environment retweeted
Impressive. Luckily we can turn @NetHack_LE / balrogai.com into an ASI benchmark in an instant. The human world record is 61 consecutive NetHack ascensions (teddit.net/r/nethack/comment…). Also curious to see if Astra can ascend in all roles (nethackwiki.com/wiki/Z-score).
9
9
90
10,488
The NetHack Learning Environment retweeted
Latest release of @NetHack_LE (v1.3.0) is now available and includes rendering the dungeon with tiles for all you convoluted people out there, seeding options for getting past the pesky full moon issue and more. Thanks to all the contributors! ⚔️

ALT Example of the NetHack Learning Environment rendering the dungeon tiles

1
4
14
2,360
Latest release of @NetHack_LE (v1.3.0) is now available and includes rendering the dungeon with tiles for all you convoluted people out there, seeding options for getting past the pesky full moon issue and more. Thanks to all the contributors! ⚔️
1
1
18
1,807
Excuse me?
So uh why is almost all of game dev done with C++ and not C?
1
2,521
1 43.6 Grok-4-Wiz-AI-Cha died in The Dungeons of Doom on level 1. Killed by a housecat.
LLMs acing math olympiads? Cute. But BALROG is where agents fight dragons (and actual Balrogs)🐉😈 And today, Grok-4 (@grok) takes the gold 🥇 Welcome to the podium, champion!
1
5
28
4,770
Is there a way to spectate games? Would be awesome if we could watch like how we do at nethack servers like alt.org and hardfought.org
1
3
129
This is an excellent idea.
1
2
45
The NetHack Learning Environment retweeted
💯 Who knew that the International Math Olympiad (IMO) is much easier than @NetHack_LE for AI.
Replying to @apples_jimmy
Meanwhile, another wall - @NetHack_LE - is still standing firm and tall.
4
6
76
8,878
The NetHack Learning Environment retweeted
Happy "@NetHack_LE is still completely unsolved" day for those of you who are celebrating it. We released The NetHack Learning Environment (arxiv.org/abs/2006.13760) on this day five years ago. Current frontier models achieve only ~1.7% progression (see balrogai.com). For a recent blog post on what makes it so hard for AI, check out @HenaffMikael's analysis: mikaelhenaff.substack.com/p/…
3
28
137
27,102
Great post by @HenaffMikael (after ascending, what an achievement!) on what makes @NetHack_LE so extremely difficult for AI (even LLMs: balrogai.com/). "While NetHack is complex in comparison to other RL benchmarks, it still contains only a tiny fraction of the complexity of the real world (its source code is 4.2MB, which provides an upper bound on its Kolmogorov complexity). As long as we can’t reliably solve this game for which we can easily collect lifetimes worth of data, have access to detailed textual resources (and even the underlying source code), and large-scale datasets of human gameplay, I think AGI remains a ways off."
A couple bits of news: 1. Happy to share my first (human) NetHack ascension-next step is RL agents :) 2. I wrote a post discussing some @NetHack_LE challenges & how they map to open problems in RL & agentic AI. Still the best RL benchmark imo. mikaelhenaff.substack.com/p/…
1
12
36
6,246
Probably, since you're presumably been able to make progress in the real world which is more complex than NetHack.
1
1
99
The NetHack Learning Environment retweeted
A couple bits of news: 1. Happy to share my first (human) NetHack ascension-next step is RL agents :) 2. I wrote a post discussing some @NetHack_LE challenges & how they map to open problems in RL & agentic AI. Still the best RL benchmark imo. mikaelhenaff.substack.com/p/…
5
13
62
11,687
The NetHack Learning Environment retweeted
Happy to announce the latest release of @NetHack_LE (version 1.2.0). You can now use the seed function to make the dungeon layout reproducible across training episodes. The in-level interaction and combat is still randomly determined and doesn't impact lower level layouts.
1
3
28
7,758
The NetHack Learning Environment retweeted
⚔️ MiniHack Updates! ⚔️ 1️⃣ MiniHack 1.0.0 is here! Following popular demand, it now supports the new Gymnasium API and is built on NLE 1.1.0. Huge thanks to @Stephen_Oman (maintainer of @NetHack_LE ) for his outstanding contribution! 🙌
3
14
65
5,064
The NetHack Learning Environment retweeted
Can AI agents adapt zero-shot, to complex multi-step language instructions in open-ended environments? We present MaestroMotif, a method for AI-assisted skill design that produces highly capable and steerable hierarchical agents. To the best of our knowledge, it is the first method that, without expert labeled datasets, solves compositional tasks requiring hundreds of steps for completion. All the modules within MaestroMotif are learned from interaction: from the highest level of planning to the lowest-level of sensorimotor control. On the open-ended domain of NetHack, it surpasses existing approaches, including those that are fine-tuned specifically for each task. At the heart of MaestroMotif is the idea that decomposing a task into subtasks significantly helps decision making. MaestroMotif leverages an agent designer's intuition about a domain to identify important skills and describe them in natural language. These short descriptions then get converted into adaptable hierarchical agents through AI feedback and in-context learning. Our paper was recently published at ICLR 2025 and we open-source the whole project including the code, prompts and pre-trained models. Paper: arxiv.org/abs/2412.08542 Code: github.com/mklissa/maestromo… NotebookLM Podcast: bit.ly/4jLi6mo This work was done with the amazing @HenaffMikael, @robertarail, @shagunsodhani, Pascal Vincent, @yayitsamyzhang, @pierrelux, Doina Precup, with equal supervision by @MarlosCMachado and @proceduralia. Take a look at the following thread:
6
51
199
80,317
I quite like the idea using games to evaluate LLMs against each other, instead of fixed evals. Playing against another intelligent entity self-balances and adapts difficulty, so each eval (/environment) is leveraged a lot more. There's some early attempts around. Exciting area.
Replying to @karpathy
Perfect timing, we are just about to publish TextArena. A collection of 57 text-based games (30 in the first release) including single-player, two-player and multi-player games. We tried keeping the interface similar to OpenAI gym, made it very easy to add new games, and created an online leaderboard (you can let your model compete online against other models and humans). There are still some kinks to fix up, but we are actively looking for collaborators :) If you are interested check out textarena.ai/, DM me or send an email to guertlerlo@cfar.a-star.edu.sg Next up, the plan is to use R1 style training to create a model with super-human soft-skills (i.e. theory of mind, persuasion, deception etc.)
250
403
5,794
979,955
Cool idea!
22
7
584
41,740
The NetHack Learning Environment retweeted
💯 For me this is NetHack (see @NetHack_LE and balrogai.com/). I am still holding my breath.
It can be hard to “feel the AGI” until you see an AI surpass top humans in a domain you care deeply about. Competitive coders will feel it within a couple years. Paul is early but I think writers will feel it too. Everyone will have their Lee Sedol moment at a different time.
3
1
19
4,595
The NetHack Learning Environment retweeted
In the meantime, @NetHack_LE and balrogai.com...

ALT Dubsado Rafy GIF

So everyone, each of us now has to work hard to develop new benchmarks. Because oh boy will they be solved quickly.
1
2
11
4,678