We're excited to release PenguinHarness 0.2.13.
What infrastructure will it take to bring recursive self-improvement into industry use?
We're building PenguinHarness around that question. Our vision is an open-source workspace where everyone can build agents, put them to work, and keep improving how they operate.
If you're new here, PenguinHarness is an Agent Harness with a CLI and Web UI, support for 1,000+ models, and a local-first design. You can configure and run agents without writing code.
The core design is that an agent's behavior lives in readable, editable files that can be versioned. Its role, operating procedures, Skills, and runtime settings form its Agent State. Sessions and Traces record what happens when it runs.
One agent can use those records to evaluate another, revise its State, and test the revision against the same benchmark. The model weights stay unchanged. Our focus is the improvement loop around the model, and the infrastructure needed to make that loop useful in everyday work.
0.2.13 develops three parts of that infrastructure: measurement, organization, and reliable execution.
1. Measurement: a way to tell whether a change helped.
The Evaluation Center now puts Benchmarks at the Project level, so multiple agents can be evaluated against the same benchmark. Results are grouped by agent, model, and thinking level, making revisions easier to compare.
"Ask AI" brings the current evaluation into a pre-filled conversation for analysis and follow-up, connecting the evaluation to a discussion of what to change next.
Comparable results give the improvement process a basis for deciding what to keep. They also give people a way to inspect the evidence behind that decision.
2. Organization: clear responsibilities and human decisions.
Company mode is now in beta, disabled by default. Give a Project a mission, and it starts with a CEO. The mode adds reporting lines, a calendar, a ticket board, channels that require an @-mention to trigger a response, and per-agent budgets.
Hiring, budget decisions, and closing P0/P1 tickets require your approval. Organization and operating state live in files on disk, with SQLite used only as a cache.
As more agents take part in the work, people need to define responsibilities, control spending, and decide which actions require approval. Company mode is our starting point for making those arrangements explicit.
3. Reliable execution: keeping the improvement process running.
An improvement loop involves running tasks, inspecting results, changing files, and running again. This release addresses several of the interruptions and constraints along that path:
- File management now supports renaming, moving, and deleting files, with larger text previews and smoother browsing. The files behind an agent become easier to inspect and organize.
- Desktop tray support keeps the server and background tasks running when you close the window on Windows, macOS, and Linux.
- Background execution lets shell commands and subagent runs move to the background after 10 seconds. The current turn can continue without stopping the process.
- Gateway fixes address task interruptions caused by image-message placement, DeepSeek compatibility, and parallel tool calls.
- Shell sandboxing defines constraints on filesystem and network access. If the backend cannot enforce the requested constraints, the command does not start.
- Model integration updates move 43 OpenRouter presets to the Responses API, add Atria-Dawn-Preview and the free dots-3-note-preview, and fix protocol detection across URL formats.
These changes support the repeated execution, file access, and model calls that the improvement process depends on.
We hope our users can own this process: understand how their agents work, decide what they should do, and use the results to make the next revision better.
So now, I highly recommend deploying your own evolvable agent platform; it will lead you to discover the charm of RSI:
penguin.ooo/