I've been experimenting with lights-out, fully automated software development for half a year. I'm pretty sure Shopify's River and Block's Builderbot are really solid but there's not a lot of info available about them. So I'll share the patterns that I've arrived at with my FactoryX system. Having a picture like this half a year ago would have saved me a lot of time.
1. Human software engineering is well represented in a kanban board. I haven't found that carries over to agentic engineering. There are really two touch points where I interact: (a) when I describe what I want out of a deliverable, and (b) after agentic QA when I review the work. As long as work is flowing efficiently I don't really need to see tickets bouncing around a board. So in
FactoryX the interface is centered on triggering work and an optional queue with build review and code review.
2. I used a software factory, modeled after a team, as the top level organizing abstraction. It has shared memory, secrets for LLMs/git/slack/..., agentic QA gates, etc. If any of those diverge, it's easy to spin up another Factory for it. I experimented with both tighter (projects — see OpenAI's Symphony) and looser (sets of factories) groupings. A single Factory is a flexible and simple abstraction.
3. The system enforces a simple documentation flow with markdown files stored alongside code, organized by work order. Requirements.md informs design.md, then human feedback from multiple channels (discord, teams, admin ui) gets folded into feedback.md. The repo itself carries the project documentation memory and state, instead of a separate datastore.
4. For large deliverables the system uses a ralph loop to just-in-time synthesize a ticket DAG, where individual tickets get executed using a /goal prompt.
5. Inspired by OpenClaw: I want multiple channels for interacting with the system, sandboxed agents doing the work, and a gateway for protecting secrets so the agents never hold them.
6. I want to mix and match harnesses (codex / opencode / claude / pi) and models (DeepSeek, Qwen, Claude, GPT…) depending on the job. A simple way to do this is to implement a gateway supporting all LLMs I want to use, and then build a single agent container containing all the harnesses.
7. I found it really helpful to implement DORA metrics and LLM telemetry (prefill / gen etc) to understand the effectiveness of the system.
8. Biggest ROI is investing cycles into automating agentic QA so that when the work reaches human review it's generally as high quality as possible.
Full pipeline diagram below. If you're building something similar or have found really good open source or commercial tools, I'd love to compare patterns. DMs are open.