Agents, MCP & Developer Experience at IBM โšก ๐Ÿฅ‘ Tech Author, Speaker, Entrepreneur ๐Ÿš€ Sharing opinions (โœ‰๏ธ hello@hackteam.io)

Santa Clara, CA
๐Ÿพ SOFTWARE DEVELOPER OF THE YEAR Iโ€™ve won a @hackernoon Noonies Award for my contributions to the developer community Thanks everyone for voting ๐Ÿ™Œ ๐Ÿ™Œ
13
7
145
Came back to the US from a trip to Europe to visit family, and magically my ambition and inspiration already tripled just by waking up here
4
11
815
My favorite question during an AI app demo: what happens if I change my mind halfway through?
1
444
I'm already using cheaper subagents, but the model planning and coordinating their work still burns through my quota. Give me usage broken down by task and agent. I want to see which part of the workflow spent the budget before I start swapping models.
1
416
freedom is driving your car without navigation on
1
3
469
Roy Derks ๐Ÿš€ retweeted
Stop caring so much about model releases. Focus on the harness, and improving the environment your agent operates in. You'll find yourself far ahead of the curve.
141
147
3,961
167,229
A useful prompt for an agent bill: 'Compare these two weeks of traces and invoices. Which tasks got more expensive? Did calls, retries or output length increase? Separate price changes from usage changes. Show the evidence.' That's a report I could actually act on.
1
396
Before an agent writes an API integration, give it this job: 'Find a real request and response, the auth flow, rate limits and error cases. List what the docs leave unclear.' I want the missing pieces visible before we have a whole app built around assumptions.
1
3
441
For a daily agent, I'd specify what happens after a missed run. Skip it? Catch up once? Replay every missed day? For my draft workflow, a current batch makes sense. Three backdated batches would just give me more material to sort through.
424
Prompt before connecting an agent to another service: "Show which account, workspace and environment the tool is using. Match stable IDs where available. If two destinations have the same name, stop and resolve that ambiguity." Names alone are a weak check.
2
401
GitHub's Rust rewrite gave SDK users an in-process runtime instead of requiring a separate Node process. That's a concrete reason to reconsider a language: changing how the software can be embedded, beyond making the same code run faster.
1
270
I want compiler feedback in the coding-agent loop. But if the agent fixes a TypeScript error by adding 'as any', we may have just removed the feedback. I'd review new casts and suppressions as carefully as the code they make compile.
2
1
432
Iโ€™m very bullish on the harness as a product, similar performance as Jev could be done with a great harness and a combination of frontier and cheap/fast OSS models
1
418
Roy Derks ๐Ÿš€ retweeted
jev is fun to play with but you have to be careful shipping it to prod, over the past few days we accidentally racked up $3.55 in usage
104
140
7,706
244,520
Roy Derks ๐Ÿš€ retweeted
Replying to @anshnanda
we built muse from scratch, but it is definitely heavily inspired as a product by openclaw. After I used openclaw in january I bought hundreds of mac minis for the MSL team and lots of us fell in love with using openclaw (and other personal agents). @steipete is a genius and his harness was pioneering from the jump. I think a lot of people were inspired by it. our goal with muse was to build something like openclaw that we could make safe and secure and easy to use and scale to billions of people.
108
243
5,124
1,223,376
Let's rename every "system one model pattern" model to decision model going forward
Jev is awesome, but its name "System One model" is not. I asked Jev what it thought it should be called, using excerpts from typesafe and OpenRouter's release posts. Here's what Jev wants to be called:
455
this is a great way to get access to the system one model pattern in the browser
Jev is the most exciting AI release in months. So I built the open, in-browser version ๐ŸŽ‰ open-jev: System One style decisions running 100% on your device. Text in, typed answers out, one forward pass. Nothing is generated, so it can't hallucinate.
3
712
Roy Derks ๐Ÿš€ retweeted
When a multi-agent demo finds more bugs, check whether it also searched more code and spent more tokens. Compare the same scope and budget before crediting the architecture. A bigger search can produce a bigger result without being more efficient.
1
3
493
This is a great stack, cost effective and performant (Pretty close to my current stack)
My model + harness stack these days: 1. Kimi K3 is the best OSS coding model. GLM-5.3 is next. 2. DeepSeek V4.1 Flash is my pick for easy/moderate tasks: insanely fast and cheap. 3. I reach for Astra for complex tasks. 4. Codex works better with OSS models. Claude Code can make the same model cost 2x more.
1
462
Developers like faster & cheaper, one of the reasons why open weight models and Jev are so hyped now
454
Time to ditch CLAUDE.md
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
4
521