🚀 Introducing rLLM v0.2 - train arbitrary agentic programs with RL, with minimal code changes.
Most RL training systems adopt the agent-environment abstraction. But what about complex workflows? Think solver-critique pairs collaborating, or planner agents orchestrating multiple workers. These are hard to express with traditional RL abstractions.
v0.2 introduces AgentWorkflowTrainer, built on a simple insight: any agentic flow is just a Python program orchestrating LLM calls, so we made ANY Python program trainable.
Researchers and developers can now quickly prototype new ideas or transform their production agentic systems into trainable flows with minimal changes.
rLLM now uses official
@verl_project ==0.5.0 as our backend (no more custom verl forks!). Just define your agentic workflow or multi-agent system as a Python program and hit train, and rLLM will handle the rest.
Since release, rLLM has been adopted to power RL training of world-class agents like
@Ali_TongyiLab's DeepResearcher.
With this new release, we're working towards building the RL application stack for next-gen agentic AI - where entire systems learn and evolve together, not just individual components in isolation.
📖 Blog post:
rllm-project.com/post.html?p…
👨💻 GitHub:
github.com/rllm-org/rllm
What agentic program will you train first? 👀