Can web agents learn through imagination?
DynaWeb: Model-Based Reinforcement Learning of Web Agents has been accepted to EMNLP 2026 Main Conference.
Training web agents with online RL means costly, risky interaction with the live web. DynaWeb learns a Web World Model instead, letting agents dream web interactions and train on imagined rollouts.
- Learns naturalistic web transitions from agent actions
- Generates imagined trajectories in a synthetic web environment
- Interleaves imagined rollouts with real expert trajectories
- Consistent gains for open-source agents on WebArena and WebVoyager
Imagination as a scalable path toward online agentic RL.
This is part of our ongoing work at Gradient on distributed and efficient RL: training capable agents without the cost, latency, and risk of centralizing them.
Paper: arxiv.org/abs/2601.22149
Code: github.com/jadeleiyu/MBRL-Ag…
Agents are impressive. Multi-agent systems are where things get interesting.
As AI systems evolve from single agents into networks of specialists, a new question follows:
How do agents decide who does what?
That’s a problem we’ve been exploring with Symphony.
Symphony brought multi-agent collaboration across heterogeneous environments. To take that further, we released Symphony-Coord, where coordination and specialization adapt through task context, performance, and feedback.
github.com/GradientHQ/sympho…
It's been rough lately, but it was all worth it in the end. AI Taiwan Expo was a blast, i had been able to connect with some wonderful people and learnt quite a lot.
This will be one of the core memory of 2026 for sure.
相信自己,相信您的理想。
Special thanks to Dexter, btw^^!
A self-evolving agent + a 428B model + 3 Macs = ?
Your own AI lab.
We ran @MiniMax_AI M3 locally with @tryParallax, right on our desk.
Then @GA_agent_ai took over to create a 5-stock portfolio and write it to disk.
No cloud. No API bills. Nothing left the machine.
Wild to see a ~3K-line agent drive all this with a 400B+ model on local hardware.
Thanks to the GenericAgent and MiniMax teams for making local AI feel real.
DeepSeek V4 flash is on par with GPT 5.4 (high), the best part is that it’s much more affordable at scale:
GPT 5.4 pro vs DeepSeek V4 flash:
Input: $30/M vs $0.14/M (214x cost difference)
Output: $180/M vs $0.28/M (643x cost difference)
Both at a million context, DeepSeek V4 Flash is really a bargain for intelligence.
🔬 lab works & stuff
come by to see some of the recent research works of the team.
📍 location: builder hub
🗓️ time: april 25th 1pm UTC
./ 🥼 coat on @Gradient_HQ
Awesome to see @tryParallax’s distributed framework for heterogeneous machines being implemented and serving up inferences!
Build and customize your own clusters for AI like never before 🤖
./ LFG @Gradient_HQ
🤖 Gradient Live Knowledge Trivia!
Come by to join us on a 20 question trivia! See where you stack up for knowledge among Grads!
📍 Location: Builder Hub
🗓️ Time: April 11th 1PM UTC
./ @Gradient_HQ memory engine on 🧠
Our cofounder @0xEricYang sat down with @yacinelearning to walk through Echo-2’s distributed RL architecture.
Dive in to learn about async RL with distributed infra, and how we are scaling this for businesses to win in the agentic era.
for those interested in distributed reinforcement learning I just finished a ~1h tutorial on the echo2 framework by @Gradient_HQ
we check:
- how to do async RL
- infra split between rollout workers and centralized learner
- interview with gradient cofounder eric yang himself!
for those interested in distributed reinforcement learning I just finished a ~1h tutorial on the echo2 framework by @Gradient_HQ
we check:
- how to do async RL
- infra split between rollout workers and centralized learner
- interview with gradient cofounder eric yang himself!
china's AI giants just launched a price war over AI coding plans.
zhipu, minimax, kimi, alibaba, bytedance, tencent all competing for the same developer wallet.
entry price: ¥7.9/mo (~$1.10) for the first month.
We're expanding pre-train and model size, it's time to explore post training.
like how we, human learn, experience, we learn as we do , on the go, everyday!
Benchmarks that test what models have memorized are saturating fast. ARC-AGI-3 is asking a harder question: can AI actually learn something new on the fly?
One direction we've been exploring: multi-agent orchestration. In our study, coordinating four frontier LLMs across multiple turns consistently matched or outperformed the strongest single model, even on tasks none of them could solve alone.
The gap between "best single model" and "best coordination of models" is where a lot of the real progress is hiding.
More on our multi-turn, multi-agent orchestration study: arxiv.org/abs/2509.23537
Our GTC takeaway is clear: NVIDIA is betting hard on open.
- NemoClaw turns OpenClaw into enterprise infrastructure.
- Nemotron 4 will be open-sourced.
- Nemotron Coalition puts eight labs on a shared open frontier model.
This is what we've been building toward. Open infrastructure for open intelligence is the direction the biggest AI companies are taking.