Recursive Self-Improvement through Multi-agent RL Post-training and Unsupervised Environment Design (UED)… but it actually works!
Delighted to finally release this paper, which trains a single LLM to act as both an Environment Designer to build new multi-turn RL training environments (using the Gym step()/reset() API), and a Reasoning Agent that learns to solve them. Resurrecting ideas from our work on UED, the Designer is trained to maximize a proxy for the Agent’s regret, computed using privileged hints.
Continuous self-improvement needs an ever-expanding supply of training environments (goals).
SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️