๐พ SOFTWARE DEVELOPER OF THE YEAR
Iโve won a @hackernoon Noonies Award for my contributions to the developer community
Thanks everyone for voting ๐ ๐
I'm already using cheaper subagents, but the model planning and coordinating their work still burns through my quota.
Give me usage broken down by task and agent. I want to see which part of the workflow spent the budget before I start swapping models.
Stop caring so much about model releases.
Focus on the harness, and improving the environment your agent operates in.
You'll find yourself far ahead of the curve.
A useful prompt for an agent bill:
'Compare these two weeks of traces and invoices. Which tasks got more expensive? Did calls, retries or output length increase? Separate price changes from usage changes. Show the evidence.'
That's a report I could actually act on.
Before an agent writes an API integration, give it this job:
'Find a real request and response, the auth flow, rate limits and error cases. List what the docs leave unclear.'
I want the missing pieces visible before we have a whole app built around assumptions.
For a daily agent, I'd specify what happens after a missed run.
Skip it? Catch up once? Replay every missed day?
For my draft workflow, a current batch makes sense. Three backdated batches would just give me more material to sort through.
Prompt before connecting an agent to another service:
"Show which account, workspace and environment the tool is using. Match stable IDs where available. If two destinations have the same name, stop and resolve that ambiguity."
Names alone are a weak check.
GitHub's Rust rewrite gave SDK users an in-process runtime instead of requiring a separate Node process. That's a concrete reason to reconsider a language: changing how the software can be embedded, beyond making the same code run faster.
I want compiler feedback in the coding-agent loop. But if the agent fixes a TypeScript error by adding 'as any', we may have just removed the feedback. I'd review new casts and suppressions as carefully as the code they make compile.
Iโm very bullish on the harness as a product, similar performance as Jev could be done with a great harness and a combination of frontier and cheap/fast OSS models
we built muse from scratch, but it is definitely heavily inspired as a product by openclaw. After I used openclaw in january I bought hundreds of mac minis for the MSL team and lots of us fell in love with using openclaw (and other personal agents). @steipete is a genius and his harness was pioneering from the jump. I think a lot of people were inspired by it. our goal with muse was to build something like openclaw that we could make safe and secure and easy to use and scale to billions of people.
Jev is awesome, but its name "System One model" is not. I asked Jev what it thought it should be called, using excerpts from typesafe and OpenRouter's release posts. Here's what Jev wants to be called:
Jev is the most exciting AI release in months. So I built the open, in-browser version ๐
open-jev: System One style decisions running 100% on your device. Text in, typed answers out, one forward pass. Nothing is generated, so it can't hallucinate.
When a multi-agent demo finds more bugs, check whether it also searched more code and spent more tokens. Compare the same scope and budget before crediting the architecture. A bigger search can produce a bigger result without being more efficient.
My model + harness stack these days:
1. Kimi K3 is the best OSS coding model. GLM-5.3 is next.
2. DeepSeek V4.1 Flash is my pick for easy/moderate tasks: insanely fast and cheap.
3. I reach for Astra for complex tasks.
4. Codex works better with OSS models. Claude Code can make the same model cost 2x more.
We're adding support for AGENTS.md to Claude Code.
Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md.
You can toggle this behavior in /config.