The least productive thing I can imagine doing is watching an agent work.
But I still do it, because I'm never really sure if the agent understood what I wanted it to do. So I end up reading the tool calls and watching files change.
I want to agree on what it's gonna build, leave it alone, and review what comes back. But an all-text plan that's several pages long just feels like homework, and it's easier to let the agent run, even if it wastes tokens, just so I can see the result.
What I need is a better way to verify the plan and review the result, just at a glance. Something where I can see the layout or the API flow and point to the part that's wrong.
Here's what I've been doing instead lately.
builder.io/blog/stop-watchin…
GPT-6 and Opus 5.5 make highly autonomous work cheaper than ever
That makes now a better time than ever to build an autonomous software factory
Here's a deep dive on how we now automated 90% of our development with self-governing agent factories
Steal our /factory skills:
🚨 Huge update: YouTube just announced video AB testing, thumbnail generation & more.
I’m here in the YouTube offices in New York… and they’ve just made their biggest update in years.
Here’s the 3 changes I am most excited about + my take as a strategist running some of the biggest channels on youtube:
1. Video AB testing
YouTube is going to allow you to upload three different versions of a video now (starting with shorts and then rolling out to more) and test them over a period to see which gets the best watch time.
My take: I am so excited by this and it feels like something we’ve all wanted for years. How often are you stuck between two different intros, or two different ways of ending a video?
Now having the ability to get real data back is insane. I am going to heavily use this with intros.
2. Dynamic thumbnail testing for different audiences
YouTube will now actually dynamically test different thumbnails with different audience segments. Allowing you to have different packaging for your returning viewers vs casual/new.
My take: This is very interesting. Current AB testing for thumbnails essentially helps you find the best thumbnail for the biggest majority of your viewers… but that isn’t the best thumbnail for every segment.
A core viewer might click on something very different to a brand new viewer. I think this will allow creators to print more views and test new styles that previously might have never won in the test.
3. Thumbnail generation in YouTube studio
Directly in your studio you will have an ‘ask studio packaging assistant’ that can analyze your video and past performance and generate thumbnails (and titles) for your video.
My take: This one’s been coming. The interesting thing for me is that when presenting this, youtube said something very interesting: they will match your style.
This is the hardest step when using most Ai thumbnail generators, so if they can pull this off… could change the game.
Really excited to see how these new features work.
Youtube shared that they will be tested and rolled out in the coming months.
🚨 Huge update: YouTube just announced video AB testing, thumbnail generation & more.
I’m here in the YouTube offices in New York… and they’ve just made their biggest update in years.
Here’s the 3 changes I am most excited about + my take as a strategist running some of the biggest channels on youtube:
1. Video AB testing
YouTube is going to allow you to upload three different versions of a video now (starting with shorts and then rolling out to more) and test them over a period to see which gets the best watch time.
My take: I am so excited by this and it feels like something we’ve all wanted for years. How often are you stuck between two different intros, or two different ways of ending a video?
Now having the ability to get real data back is insane. I am going to heavily use this with intros.
2. Dynamic thumbnail testing for different audiences
YouTube will now actually dynamically test different thumbnails with different audience segments. Allowing you to have different packaging for your returning viewers vs casual/new.
My take: This is very interesting. Current AB testing for thumbnails essentially helps you find the best thumbnail for the biggest majority of your viewers… but that isn’t the best thumbnail for every segment.
A core viewer might click on something very different to a brand new viewer. I think this will allow creators to print more views and test new styles that previously might have never won in the test.
3. Thumbnail generation in YouTube studio
Directly in your studio you will have an ‘ask studio packaging assistant’ that can analyze your video and past performance and generate thumbnails (and titles) for your video.
My take: This one’s been coming. The interesting thing for me is that when presenting this, youtube said something very interesting: they will match your style.
This is the hardest step when using most Ai thumbnail generators, so if they can pull this off… could change the game.
Really excited to see how these new features work.
Youtube shared that they will be tested and rolled out in the coming months.
Is Jev actually good at computer use, or are those videos all over twitter more fake (or highly misleading) demos?
I put it to the test and the results surprised me:
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale.
We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
I had a banked reset expiring Sep 21st. Assumed EOD 21st but it expired Sep 20th 11 59?
Wish this could be more clear. Just lost a reset. @reach_vb@thsottiaux
We're adding support for AGENTS.md to Claude Code.
Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md.
You can toggle this behavior in /config.
I think there’s a slightly underappreciated reason coding agents work so well.
It’s not just that LLMs are getting better at code.
Software engineering happens to expose an unusually agent-friendly environment:
repo → context
terminal/tools → actions
git → state + reversibility
tests/CI → feedback
PRs → human verification
Put together, this creates a tight observe → act → verify → iterate loop.
In other words, we accidentally spent decades building a pretty good agent harness around software development.
Most SaaS products don’t have this yet. The model may be equally capable, but the surrounding environment gives it much less leverage.
Wrote up the idea + what I think SaaS can learn from coding agents:
builder.io/blog/the-real-rea…
I keep getting frustrated when I ask an agent to clean up some code and get another wrapper, another helper, and all the original complexity still sitting underneath.
Then I have to figure out if anything actually got better.
We’re running into this at Builder while developing Agent Native pretty much as fast as we can. People leave feedback in Slack, and an agent factory picks it up and works on it. A lot is changing, all the time.
So validation is a big deal for us. In one cleanup experiment, an agent removed a lint warning, passed typechecking, and passed all 23 tests in the focused suite. The code still used unvalidated input.
Everything passed, but we still hadn’t checked the behavior we actually needed.
I wrote up what we’re learning about asking agents to clean up code, figuring out what needs to change, and checking whether it actually got better.
builder.io/blog/de-slop-ai-g…
I used to work at Grafana.
The dashboard-only era is ending. That's why we built Agent Native Analytics.
A chatbot that generates a plausible chart and forgets it when the tab closes isn’t AI analytics. It’s a demo.
The future is durable work shared by agents and humans.
GPT-6 Sol is Coming This week
GPT-6 Sol Will be released On Thursday
This Week will be a Big Week
Community note
The post from the unofficial @Codexresets_ account falsely claims an imminent GPT-6 Sol release this week. OpenAI has announced no such model; its current flagship is GPT-6 Astra (launched Sept 3) while Sol is part of the earlier GPT-5.6 family (launched July).
openai.com/index/gpt-6-as…openai.com/index/gpt-5-6/x.com/openai/status/…