The least productive thing I can imagine doing is watching an agent work. But I still do it, because I'm never really sure if the agent understood what I wanted it to do. So I end up reading the tool calls and watching files change. I want to agree on what it's gonna build, leave it alone, and review what comes back. But an all-text plan that's several pages long just feels like homework, and it's easier to let the agent run, even if it wastes tokens, just so I can see the result. What I need is a better way to verify the plan and review the result, just at a glance. Something where I can see the layout or the API flow and point to the part that's wrong. Here's what I've been doing instead lately. builder.io/blog/stop-watchin…
5
15
2,071
GPT-6 and Opus 5.5 make highly autonomous work cheaper than ever That makes now a better time than ever to build an autonomous software factory Here's a deep dive on how we now automated 90% of our development with self-governing agent factories Steal our /factory skills:
10
33
336
21,662
Already have a hard time editing one video. Now need to try variants Great feature tho!
🚨 Huge update: YouTube just announced video AB testing, thumbnail generation & more. I’m here in the YouTube offices in New York… and they’ve just made their biggest update in years. Here’s the 3 changes I am most excited about + my take as a strategist running some of the biggest channels on youtube: 1. Video AB testing YouTube is going to allow you to upload three different versions of a video now (starting with shorts and then rolling out to more) and test them over a period to see which gets the best watch time. My take: I am so excited by this and it feels like something we’ve all wanted for years. How often are you stuck between two different intros, or two different ways of ending a video? Now having the ability to get real data back is insane. I am going to heavily use this with intros. 2. Dynamic thumbnail testing for different audiences YouTube will now actually dynamically test different thumbnails with different audience segments. Allowing you to have different packaging for your returning viewers vs casual/new. My take: This is very interesting. Current AB testing for thumbnails essentially helps you find the best thumbnail for the biggest majority of your viewers… but that isn’t the best thumbnail for every segment. A core viewer might click on something very different to a brand new viewer. I think this will allow creators to print more views and test new styles that previously might have never won in the test. 3. Thumbnail generation in YouTube studio Directly in your studio you will have an ‘ask studio packaging assistant’ that can analyze your video and past performance and generate thumbnails (and titles) for your video. My take: This one’s been coming. The interesting thing for me is that when presenting this, youtube said something very interesting: they will match your style. This is the hardest step when using most Ai thumbnail generators, so if they can pull this off… could change the game. Really excited to see how these new features work. Youtube shared that they will be tested and rolled out in the coming months.
1
526
Vishwas retweeted
🚨 Huge update: YouTube just announced video AB testing, thumbnail generation & more. I’m here in the YouTube offices in New York… and they’ve just made their biggest update in years. Here’s the 3 changes I am most excited about + my take as a strategist running some of the biggest channels on youtube: 1. Video AB testing YouTube is going to allow you to upload three different versions of a video now (starting with shorts and then rolling out to more) and test them over a period to see which gets the best watch time. My take: I am so excited by this and it feels like something we’ve all wanted for years. How often are you stuck between two different intros, or two different ways of ending a video? Now having the ability to get real data back is insane. I am going to heavily use this with intros. 2. Dynamic thumbnail testing for different audiences YouTube will now actually dynamically test different thumbnails with different audience segments. Allowing you to have different packaging for your returning viewers vs casual/new. My take: This is very interesting. Current AB testing for thumbnails essentially helps you find the best thumbnail for the biggest majority of your viewers… but that isn’t the best thumbnail for every segment. A core viewer might click on something very different to a brand new viewer. I think this will allow creators to print more views and test new styles that previously might have never won in the test. 3. Thumbnail generation in YouTube studio Directly in your studio you will have an ‘ask studio packaging assistant’ that can analyze your video and past performance and generate thumbnails (and titles) for your video. My take: This one’s been coming. The interesting thing for me is that when presenting this, youtube said something very interesting: they will match your style. This is the hardest step when using most Ai thumbnail generators, so if they can pull this off… could change the game. Really excited to see how these new features work. Youtube shared that they will be tested and rolled out in the coming months.
80
62
1,199
879,450
Is Jev actually good at computer use, or are those videos all over twitter more fake (or highly misleading) demos? I put it to the test and the results surprised me:
13
20
204
12,470
RIP. Was hoping it wouldn't change for F1
A new look for the McLaren logo 👀
1
479
Vishwas retweeted
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
2,234
5,298
53,492
9,989,881
Great improvements!
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
4
599
sitting at 1% please reset asap. thx
Replying to @My_Ai_Bi
3am on a tuesday
1
617
I had a banked reset expiring Sep 21st. Assumed EOD 21st but it expired Sep 20th 11 59? Wish this could be more clear. Just lost a reset. @reach_vb @thsottiaux
1
1
634
My best performing video in the first 24 hours lets goo
1
10
589
Here’s my beginner-friendly guide to Jev: what it is and how to use it. Getting back into the habit of making videos, so feedback is welcome!
1
12
1,040
PSA
Replying to @udiWertheimer
OK fine. But it’s also still coming in Tuesday
1
620
Jev is awesome but for the love of god please STOP posting fake demos
213
254
3,207
303,137
Finally
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
6
801
Vishwas retweeted
I think there’s a slightly underappreciated reason coding agents work so well. It’s not just that LLMs are getting better at code. Software engineering happens to expose an unusually agent-friendly environment: repo → context terminal/tools → actions git → state + reversibility tests/CI → feedback PRs → human verification Put together, this creates a tight observe → act → verify → iterate loop. In other words, we accidentally spent decades building a pretty good agent harness around software development. Most SaaS products don’t have this yet. The model may be equally capable, but the surrounding environment gives it much less leverage. Wrote up the idea + what I think SaaS can learn from coding agents: builder.io/blog/the-real-rea…
21
5
52
15,633
I keep getting frustrated when I ask an agent to clean up some code and get another wrapper, another helper, and all the original complexity still sitting underneath. Then I have to figure out if anything actually got better. We’re running into this at Builder while developing Agent Native pretty much as fast as we can. People leave feedback in Slack, and an agent factory picks it up and works on it. A lot is changing, all the time. So validation is a big deal for us. In one cleanup experiment, an agent removed a lint warning, passed typechecking, and passed all 23 tests in the focused suite. The code still used unvalidated input. Everything passed, but we still hadn’t checked the behavior we actually needed. I wrote up what we’re learning about asking agents to clean up code, figuring out what needs to change, and checking whether it actually got better. builder.io/blog/de-slop-ai-g…
3
6
1,135
Vishwas retweeted
I used to work at Grafana. The dashboard-only era is ending. That's why we built Agent Native Analytics. A chatbot that generates a plausible chart and forgets it when the tab closes isn’t AI analytics. It’s a demo. The future is durable work shared by agents and humans.
2
3
18
2,433
Vishwas retweeted
GPT-6 Sol is Coming This week GPT-6 Sol Will be released On Thursday This Week will be a Big Week
Community note
The post from the unofficial @Codexresets_ account falsely claims an imminent GPT-6 Sol release this week. OpenAI has announced no such model; its current flagship is GPT-6 Astra (launched Sept 3) while Sol is part of the earlier GPT-5.6 family (launched July). openai.com/index/gpt-6-as… openai.com/index/gpt-5-6/ x.com/openai/status/…
212
269
5,883
512,721