Research Engineer @ Google Deepmind | PhD Student @ University of Toronto.

Ryan Faulkner retweeted
Recursive Self-Improvement through Multi-agent RL Post-training and Unsupervised Environment Design (UED)… but it actually works! Delighted to finally release this paper, which trains a single LLM to act as both an Environment Designer to build new multi-turn RL training environments (using the Gym step()/reset() API), and a Reasoning Agent that learns to solve them. Resurrecting ideas from our work on UED, the Designer is trained to maximize a proxy for the Agent’s regret, computed using privileged hints.
Continuous self-improvement needs an ever-expanding supply of training environments (goals). SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
13
56
426
57,037
Excited to share our latest research and open-source codebase: AgentElect! 🚀 We investigate how AI systems can leverage elections to navigate resource-based social dilemmas and drive multi-agent cooperation in common pool resource problems. 📄 Read the paper: arxiv.org/abs/2604.11721💻 Explore the code: github.com/rfaulkner/GovSimE… #MultiAgentSystems #AI #MachineLearning #AIResearch #OpenSource
1
3
41