The NetHack Learning Environment (@NetHack_LE) has been around since 2020. Since 2024, balrogai.com is evaluating frontier LLM performance on it. We are slowly seeing progress on it, but it still seems far from being solved.

Sep 21, 2026 · 4:02 PM UTC

7
6
110
59,886
Sort replies: Relevant Recent Liked
Replying to @_rockt @NetHack_LE
I love nethack!
4
3,628
Replying to @_rockt @NetHack_LE
nethack still humbling the frontier models lol
2
2,910
Replying to @_rockt @NetHack_LE
yeah nethack's the one eval reasoning models can't brute force. it's all long-horizon memory over thousands of steps
2
880
Replying to @_rockt @NetHack_LE
My vision of good reasoning AI is one where the environment is faced for the first time (no pretraining) and it is capable of figuring out strategies based on possible actions. I suspect we will see improvement on current systems if these evaluations fall into their radar
2
1,581
Replying to @_rockt @NetHack_LE
NetHack humbles models
Introducing Halo, the best framework for post-training of open-source models. Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format. Star us on GitHub: github.com/whitecircle/halo
2
648