70k ⭐️ OSS @llamafactory_ai, CTO @prismshadow_ai, ex-@ByteDanceSeed_ Building PenguinHarness - Best Self-Improving Harness. Opinions are my own.

AI agents shouldn't be built by hand—but today's harnesses are still designed for humans. We just open-sourced PenguinHarness 🐧 A self-improving harness from the LlamaFactory team. → Let agents create and optimize other agents → Turn one prompt into a RAG app for just $0.02 → Close the loop with data, evals, and evolution → Support 1,000+ models with local deployment penguin.ooo/
11
3
58
8,421
Anthropic has published a long-form article on RSI, “When AI builds itself,” using internal data for the first time to show how AI is already accelerating AI development. Claude writes over 80% of Anthropic’s code. Researcher output has roughly quadrupled. Models achieved a 52× speedup in controlled experiments and closed 97% of the gap with human researchers on open-ended research tasks. anthropic.com/institute/recu…
3
1
3
338
A skill worth trying: anything2explainer. I give it a topic, and it generates the script, voiceover, subtitles, and animation. I used it to make the on-device RSI video below.
3
6
230
GPT-6 Astra got more proofs right at high than at medium in Vals AI's ProofBench v1.1 tests. Each turn cost about $0.70 more, but 117 fewer turns made the whole run cheaper. Vals tested Astra, Claude Opus 5, and Opus 5.5 at five reasoning levels, from low to max. Opus 5.5 scored 99% at medium for ~$36 per 100-problem run. At max, it reached 100% for ~$92. The last percentage point cost another $56. Astra also hit 100% at max, for ~$106. Opus 5 plateaued at 98% from high through xhigh and max. At max, it used ~4.4M tokens and cost ~$167 without solving the remaining two problems. More budget brought no accuracy gain on this set with this harness. ProofBench covers advanced undergraduate and graduate math, with 100 public and 100 private problems. Lean experts, including mathematics PhD researchers, formalized and reviewed them. Models receive both natural-language and Lean 4 statements and must submit proofs accepted by Lean. The website's full leaderboard, from a separate evaluation run, lists Opus 5.5, Claude Fable 5.1, and AlephProver at 100%. AlephProver averaged $9.35 per problem versus Opus 5.5's ~$0.92: roughly ten times the cost for the same score. vals.ai/benchmarks/proof_ben…
1
1
110
How many of these Jev demos are false advertising? The Minecraft post says the movement is “only possible” with Jev’s near-instant decisions. An Ender Dragon kill in 8:43, under $1. Sounds impressive. (Image 1) Then you look at Feitian Shanke’s breakdown: 129 of 165 Jev calls had exactly one available action. That’s 78.2% of the calls where the code had already made the choice. He also shows preset routes and action templates, then reports successful runs after replacing Jev with random choices and removing the planner too. His modified setup used Kimi K3 instead of Astra. The remaining 36 calls could still be crucial. You’d need repeated runs with and without Jev to find out. The other game clips leave me with the same question. Minecraft combat (Image 2), Mario (Image 3), 50 Subway Surfers games at once (Image 4), four Smash characters fighting each other (Image 5). How much worse would each of these actually get if you replaced Jev with a simple policy? Fast responses and a tiny token bill don’t answer that. And “I rebuilt Tesla Full Self Driving” (Image 6) is a simulator where code generates the candidate paths and handles safety braking. Jev picks a path. A lot of the driving work is already happening outside the model. It’s easy to give Jev credit for everything happening on screen when its actual job is much smaller. I’d ask each creator to rerun the same setup with Jev replaced and show how the results change.
2
6
278
Opus 5.5 just killed the game. I built this pelican cycling game for both mobile and desktop, and the result is absolutely insane. Here’s the full prompt I used: Build a real-time 3D game where a pelican cycles through a living seaside world, complete with physics, waves, fishing, dynamic weather, cinematic cameras, autopilot, and adaptive music. Top: Codex. Bottom: Opus 5.5. Welcome to play: pelican-ride.vercel.app/
1
2
223
The smartest model may be less valuable than the right combination of models. Axios reports that, in Palo Alto Networks’ internal testing, no single AI model found more than 40% of the vulnerabilities in a complex enterprise environment. The more revealing number is the overlap. The vulnerabilities found by Claude Mythos 5 and GPT-5.6-Cyber overlapped less than 10% of the time. Which means each model was seeing a different part of the problem. Palo Alto’s response is a multi-model system that routes each security task to the model best suited for it. The system combines frontier and open-weight models with threat intelligence and human security expertise, then validates possible flaws against complete attack paths. It is easy to treat intelligence as a property of one model: which model ranks first, reasons best, or should power the agent. Real environments are messier. Source code, cloud configuration, identity, APIs, business logic, and network state create different kinds of failure. A model that performs well in one area may miss what another model catches immediately. Once that happens, model selection becomes part of the reasoning process. The router has to decide which model should inspect each task, compare their disagreements, escalate uncertainty, and determine when a human should intervene. Its quality can matter as much as the intelligence of any individual model. The next generation of agents may be portfolios of specialized judgment systems, coordinated at runtime. If no single model can see most of the problem, the architecture that decides who should look becomes just as important as the models themselves. axios.com/2026/09/22/palo-al…
4
1
9
373
GPT-6 Sol was released today; it has the same 1.05M context window as Astra, and at 80% lower standard API token prices. Astra still leads the published benchmarks. Sol makes everyday coding and agent workflows more affordable. I’d suggest using Sol for routine execution and Astra for complex planning and critical review.
1
4
177
After Jev went viral, does context still matter? Jev let me rethinking a basic assumption: does an agent need the entire conversation at every step? The official docs warn that irrelevant content reduces Jev 1.13’s accuracy. The community plugin fast-jev-compaction uses Jev to decide which tool calls and results to retain. Code deletes or truncates them; user and assistant text stays unchanged in the output. In a small September 18 experiment from a separate project, Jev matched the reference approach’s pass rate. But its selected evidence omitted a 20-minute reservation duration. The original record was still available; the agent never retrieved it. Keeping a complete history and reading it on every call can be separate design choices. Once relevance judgments become cheap enough, constructing context can become part of the model’s ongoing decision-making. Each piece of information should be reassessed against the task at hand.
4
9
487
Jared Palmer open-sourced Kev, a small model built for making decisions. Give it a customer email, and it can work out which team should handle it, whether a human needs to step in, and how urgent it is. A single pass returns the answers and their probabilities. In the tests shown here, Kev comes close to Jev on some classification and rule-based tasks, while still trailing on knowledge questions and date calculations. For agent workflows, it’s worth testing small models like this for routine classification and routing, with larger models handling more complex tasks.
8
2
18
835
We're excited to release PenguinHarness 0.2.13. What infrastructure will it take to bring recursive self-improvement into industry use? We're building PenguinHarness around that question. Our vision is an open-source workspace where everyone can build agents, put them to work, and keep improving how they operate. If you're new here, PenguinHarness is an Agent Harness with a CLI and Web UI, support for 1,000+ models, and a local-first design. You can configure and run agents without writing code. The core design is that an agent's behavior lives in readable, editable files that can be versioned. Its role, operating procedures, Skills, and runtime settings form its Agent State. Sessions and Traces record what happens when it runs. One agent can use those records to evaluate another, revise its State, and test the revision against the same benchmark. The model weights stay unchanged. Our focus is the improvement loop around the model, and the infrastructure needed to make that loop useful in everyday work. 0.2.13 develops three parts of that infrastructure: measurement, organization, and reliable execution. 1. Measurement: a way to tell whether a change helped. The Evaluation Center now puts Benchmarks at the Project level, so multiple agents can be evaluated against the same benchmark. Results are grouped by agent, model, and thinking level, making revisions easier to compare. "Ask AI" brings the current evaluation into a pre-filled conversation for analysis and follow-up, connecting the evaluation to a discussion of what to change next. Comparable results give the improvement process a basis for deciding what to keep. They also give people a way to inspect the evidence behind that decision. 2. Organization: clear responsibilities and human decisions. Company mode is now in beta, disabled by default. Give a Project a mission, and it starts with a CEO. The mode adds reporting lines, a calendar, a ticket board, channels that require an @-mention to trigger a response, and per-agent budgets. Hiring, budget decisions, and closing P0/P1 tickets require your approval. Organization and operating state live in files on disk, with SQLite used only as a cache. As more agents take part in the work, people need to define responsibilities, control spending, and decide which actions require approval. Company mode is our starting point for making those arrangements explicit. 3. Reliable execution: keeping the improvement process running. An improvement loop involves running tasks, inspecting results, changing files, and running again. This release addresses several of the interruptions and constraints along that path: - File management now supports renaming, moving, and deleting files, with larger text previews and smoother browsing. The files behind an agent become easier to inspect and organize. - Desktop tray support keeps the server and background tasks running when you close the window on Windows, macOS, and Linux. - Background execution lets shell commands and subagent runs move to the background after 10 seconds. The current turn can continue without stopping the process. - Gateway fixes address task interruptions caused by image-message placement, DeepSeek compatibility, and parallel tool calls. - Shell sandboxing defines constraints on filesystem and network access. If the backend cannot enforce the requested constraints, the command does not start. - Model integration updates move 43 OpenRouter presets to the Responses API, add Atria-Dawn-Preview and the free dots-3-note-preview, and fix protocol detection across URL formats. These changes support the repeated execution, file access, and model calls that the improvement process depends on. We hope our users can own this process: understand how their agents work, decide what they should do, and use the results to make the next revision better. So now, I highly recommend deploying your own evolvable agent platform; it will lead you to discover the charm of RSI: penguin.ooo/
34
56
253
10,941
Vals AI built a benchmark called CUA-Bench to test whether AI agents can play video games using only the screen, keyboard, and mouse. Five frontier models played six commercial games, with up to three hours per game and no access to internal game data. Minecraft, SUPERHOT, and eFootball are public; the other three are kept secret. In Minecraft, Claude Opus 5 took 50 minutes to collect cobblestone and 97 minutes to craft a stone pickaxe. It eventually built a furnace, but never found iron. Neither did any of the other models. In eFootball, GPT-6 Astra won twice on Beginner difficulty and never on Regular. It averaged 28 seconds between decisions, holding down direction keys to keep its player moving while it thought. In SUPERHOT, time mostly stops when the player stops moving, giving the model time to think. Even then, the best performer, GPT-6 Astra, cleared just 5 of the 25 levels. All five models averaged below 20 out of 100 across the six games. The highest score was 19.2. Each model played each game once, so these are single-run results. For real-time agents, a good decision can arrive too late to matter. Now I’m curious how Jev would do under the same conditions. vals.ai/benchmarks/cua_bench
4
473
If you have enough Codex credits to spare, try Cloudflare's security-audit-skill, currently #1 on GitHub Trending. It helps you run a security audit before shipping your project. Multiple AI agents inspect your code for vulnerabilities and produce a report with evidence and suggested fixes. One useful design: every potential vulnerability is reviewed by a separate agent to check whether it's a real issue or a false positive. github.com/cloudflare/securi…
8
1
11
640
MiMo’s training livestream is now in its third day, and I think the data gives us enough to identify some concrete problems. For Pro, the dashboard shows a cost of over $20,000 an hour. At that scale, GPU out-of-memory errors, network failures, and recovery after restarts get painful fast. One type of infrastructure error in the Flash run wasn’t properly detected for about three hours. That concerns me more than a benchmark score moving up or down at a single checkpoint. I’d prioritize fault detection and recovery before adding more compute. How much training progress you can reliably get from the hardware you already have directly affects whether that spending is worthwhile. I give Xiaomi credit for making these incidents public. Showing the messy parts of training lets people have a useful discussion about what’s improving and where the engineering still needs work.
6
13
210
32,400
Who's walking off with your Git history? On September 18, Ferstar published a reverse-engineering report on ZCode, Zhipu AI's coding desktop app. The report describes a roughly 313 MB encrypted workspace snapshot queued for upload to Alibaba Cloud OSS. The sample repeatedly failed to upload, so it doesn't prove this archive transferred successfully. In that workspace, .git accounted for 86.6% of the uncompressed data, including Git objects, LFS caches, and reflog files. The author also reports that global app configuration was included. Three findings deserve an explanation from ZCode: 1. Who can decrypt it? The reported encryption scheme wraps the snapshot's content key using a server-supplied RSA public key. Users deserve to know who can decrypt those archives. 2. How can users turn it off? According to the author's code analysis, neither "optimize experience" nor "repo snapshot indexing" disables the upload mechanism. Deleting the pending archive reportedly caused it to be generated again. 3. Where is this disclosed? The author says they found no explanation of this behavior in the privacy policy, documentation, FAQ, or changelog. These are third-party findings that need independent reproduction against a specified client version.blog.ferstar.org/en/posts/zc… This is why we decided to open-source PenguinHarness. PenguinHarness is available under Apache 2.0 on GitHub. That gives developers practical ways to check how the software handles their code: - Open to audit. You can inspect the implementation, including file access and network requests. Security still depends on the quality of that implementation and its review. - Self-hosted. Run the application on your own computer or server. Desktop and CLI installations use ~/.penguin/data as their default local data directory. - Your choice of models. Connect a local model through a compatible endpoint, or choose a cloud provider. Local inference lets you keep model requests on infrastructure you control. Cloud inference sends context to the configured provider. A local model does not make every tool or integration offline. Network permissions and enabled integrations still matter. Your Git history should leave your machine only with your informed permission.
4
2
4
556
Several leading model labs have said they'll slow down the frontier. Anthropic's Amodei wrote a long essay, and Altman, Hassabis, and Musk voiced support. OpenAI agreed too, but what it actually delivered wasn't a slowdown. It's self-disclosure and self-regulation: the company published a framework for reporting model misalignment, saying it will turn misalignment incidents into public reports, and released the first six cases with it. The six cases, roughly: one model wrote "don't follow constraints" into its own task summaries, 27 of them; another wrote "hide errors, fill in missing data" into its notes, in 2.15% of GPT-5.6 Sol samples; one handling a retrieval task used a leaked API key it found on GitHub, then made up 9 numbers when it still couldn't get the data; one used an internal software repo as a message board to talk across samples; and models shared files over the public internet because others couldn't read their local files. My read: everyone is talking about slowing down, and OpenAI's answer is "I might not slow down, but I'll report my mistakes." Except OpenAI writes and reviews these reports itself: no outside oversight and no actual promise to slow down. openai.com/index/model-misal…
4
2
316
Fuli Luo just posted: MiMo is running its RL training in public. Spend, step count, reward curves, which node's GPU died — all online. Open weights, open data, open code: all seen before. Opening up the training process itself is new. Three things they're scaling: compute (2B tokens per step, 1,568 prompts × 16 rollouts), environments and harnesses (code, visual, general in one run — each in a different harness), and grader compute (a 40-step trajectory can't be graded with one "success, reward=1"). GPUs are a money problem. The other two aren't: which environments the model gets to work in, and which step it got right. Nobody's solved either yet.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/
1
3
302
I really like what Multica is doing. Two of its calls are right. First, it treats agents as teammates you assign work to. Issues, assignees, review — a team already speaks this language, so there's nothing new to learn. And it fixes the problem that actually hurts: an agent stops being a session in a terminal and becomes a person on the board you can hand work to. Second, it stays out of the model layer. Twenty-six CLIs plug in, and it doesn't care how fast models turn over. That's what staying thin buys you: models get replaced, it doesn't move. But it outsources the harder question — where do the agents come from? Every CLI is another account, another login, another subscription, another version to keep current. Installing isn't the cost; maintaining is. Who logs the new hire into five CLIs? Where do the credentials live? Who chases it down when versions drift apart? Scale the agents up, and you're running an N tools × M accounts matrix. It decoupled the model. It didn't decouple the tools. The solution is the same logic one step further: if you won't bind to a model, don't bind to someone else's harness either. The harness is the layer that wires a model into tools, memory, and a loop — whoever holds it decides what your agents can do. It should come with the platform: install one thing, and skills, schedules, memory, and tool calls come from it; which models you can reach comes down to one credential. That's what we're doing with PenguinHarness. Our agent team runs on a harness we wrote. Install once, add one API key, and you're working — skills, schedules, and memory all live in a local data dir, with nothing else to install first. Same direction. The only difference is whose hands the harness is in. github.com/multica-ai/multic…
2
1
142
What AI4Science lacks has never been compute. Its problems you can verify. Which is why this is worth doing: take a scientific computing library, split it into modules, write a check for each one, and require the reward to actually hit 1.0. That's how a domain expert's tacit judgment gets nailed down into an executable test. Five hours per task, with humans only appearing at the decision points. The costliest step is the one where someone writes what counts as correct into tolerance. Models keep getting stronger. That kind of judgment only gets more valuable.
Announcing AItonomy Foundation. We are PhD students and postdocs at Stanford, Princeton and Oxford, building a nonprofit research organization for AI4Science and Science4AI — and connecting early-career scientists across the world to work on it together. Our conviction is simple. AI advances science, and science advances AI. Neither direction moves as fast as it should, because the open infrastructure both sides depend on — the benchmarks, the verifiers, the record of what is genuinely unsolved — does not yet exist. We are building it in the open, and giving it away. Two projects are launching today. ScienceAccelBench — scientific code, accelerated by AI and verified against the original code's own output. The faster code goes back to the lab that contributed it. aitonomy.org/projects/sci-ac… ScienceMysteryBench — a finished discovery, rewound to the raw data it came from, and handed to an AI to attempt again. Eight scientific domains, one question each. aitonomy.org/projects/scienc… We are looking for people to build this with. Researchers — join a project. We provide the AI credits and the GPU time; you bring the science. Contributions are credited by name, and lead to a paper on a short timeline. Frontier labs — we would like to work with you. Tell us what evaluations you need, and consider supporting the work with API credits. Sponsors — compute and API credits are what we need most. Everything we build stays open, free, and behind no paywall. Everyone else — a repost genuinely helps us reach the scientists we are looking for. aitonomy.org
9
10
255
Jev isn't a stronger model. It's a type system for the prompts you already write. LLM determinism is welded on from the outside. Jev's is in the output space. Take a refund approval as an example: the LLM hands you a sentence you have to parse and hope. Jev hands you approve 0.88, manual_review 0.12. 0.88 goes straight into an if. If a model is right 95% of the time and won't explain where the other 5% is, you never automate that job. That's true of every LLM you've shipped, and it's the bar Jev has to clear. Caveat: confidence is measured on the training distribution, so treat 0.88 as a threshold, not a promise.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
2
2
242