Most predictions I see are still way too conservative. Here's mine
531
887
8,695
2,319,506
"Programming languages and frameworks are dead." Yes, kind of. What will matter in future is runtime behavior: Performance characteristics, resource usage, failure modes, observability, debuggability, deployments, rollbacks. And legibility of all that to agents.
37
9
171
11,509
But then look at this new simd package: go.dev/blog/simd-experiment That's pretty handy for agents, isn't it? But... is it necessary? Not sure!
1
4
1,983
Okay I revise my take: They're not dead but the way in which we chose them and why will completely change and I suspect that will affect the languages/frameworks.
1
6
1,200
The days of reading transcripts up close are over.
If you have the patience to watch your agents work step-by-step, you're giving them too short a leash. So: ampcode.com/news/less-noise
5
40
4,513
Me grappling with my wife's cousin telling me that their whole engineering team got switched to Qwen.
"Just use a cheaper model" sounds like saving money. But fiber didn't just make your internet faster. It made whole new businesses possible. @thorstenball and @sqs discuss how the same is true with frontier models.
5
3
34
5,209
So, @0xBrettj is asking me which edit of a video is better and I click on the links and... ... he built a whole custom video editor running in orbs?!
2
1
23
2,438
Hot Take Thursday!
Would you prefer to save money with dial-up when your competitors are using fiber? Recorded on Hot Take Thursday, @sqs and @thorstenball are back with a new Raising An Agent. In this episode they discuss how cheaper models end up costing more, why they do not not wait for agents, running evals in Orbs, and the hot takes Thorsten left out of his X post last week. 00:00 Intro: Engineers who forget the history of technology 01:55 Hot Take Thursday & a €50-a-week token budget 05:14 Why cheap models cost more than frontier models 08:11 Token budgets & who gets priced out 14:06 Disrupt yourself with small teams & unlimited tokens 16:38 Agents go horizontal across the software lifecycle 20:56 Rollouts, bug reports & "I'm feeling Pucky" 23:53 Code review, monorepos & other walls around the agent 25:30 How to show a cautious company what agents can do 29:48 Why most people haven't seen what agents can do 31:57 Don't wait for your agent 33:38 Simple evals & agents testing agents in Orbs 38:27 Why Amp has runners too 40:47 Follow your energy & stay aligned 46:33 The hot take Thorsten left out 54:14 Do languages & frameworks still matter?
2
21
4,790
Thorsten Ball retweeted
the moat is literally just caring about your work and being proud of your output and holy shit it does not seem to apply to 99% of people
64
273
4,362
144,354
My maybe controversial opinion: If TTFT is important to you, you're doing something wrong. Don't wait for the agents, make them wait for you. Don't stare at a single transcript. Do something else. (Yes, there's a tiny number of exceptions to this.)
playtesting every harness again, i cannot believe how you folks don't lose it over the TUI taking over 50ms to start, and TTFT >2s - these companies are achieving AGI and here we're stuck with the slowpoke harnesses wtf
14
83
8,726
Time to share your Amp runners ampcode.com/news/shared-runn…
12
6
122
84,330
Increasingly convinced: people who do best with agents are those that are good at shipping things in small pieces. Someone who could ship a feature in 4 commits going to main now does better with agents than the guy who always ended up with a 6k PR on Friday afternoon.
67
28
664
29,274
Another thought I had yesterday in that conversation: These models are so much more now than text-to-code converters. Think bigger! Aim higher! Now that the models can contribute across your whole stack, vertically, it's to let them loose horizontally: Let them monitor deployments, let them debug prod, let them do ops, let them help end to end. If you think "they'll screw this up" that's on you and your codebase. Either you need to spend more tokens and build up instincts or you need to make your codebase and company processes friendlier to agents.
Unexpectedly found myself in a conversation about agents yesterday evening with someone who now has to use Qwen at work, because company wants to save money. Really struggled with expressing how much of a category change using latest frontier models is vs. using Qwen.
11
6
109
10,150
Unexpectedly found myself in a conversation about agents yesterday evening with someone who now has to use Qwen at work, because company wants to save money. Really struggled with expressing how much of a category change using latest frontier models is vs. using Qwen.
55
13
565
147,903
When you have friends around the world you realize: It's always orbin' time somewhere
1
29
3,041
Thorsten Ball retweeted
still can't get over how fun (and productive) multiplayer amp orbs are frontend changes used to be sooo slow to iterate on
4
4
45
4,349
"the best harness in the market"
Replying to @AmpCode
@AmpCode is the best harness in the market, there's no point even doing a tier list
6
1
81
9,414
Same prompt, four modes/models: - GPT-6 Sol - Opus 5.5 - High (GPT-6 Astra) - Ultra (Fable 5.1) Here's the prompt: "Look at the last two news posts we published on runner functionality: one runner is now enough, and one runner many worktrees. I now want you to write a news post that's similar but for the functionality behind the `shared-runners` feature flag. It should probably also contain a screenshot of the picker that shows two normal runners and two shared runners, like `gpu-runner` or `macos-builder` or something and some nic elooking avatars and good data. And it should succinctly explain how it works and how people can use it." Costs: - Opus 5.5: $9.54 - GPT-6 Sol: $3.10 (estimated list price, using sub) - High: $7.58 (estimated list price, using sub) - Ultra: $15 Notes: - Ultra took by far the longest, Ultra is the only one that didn't create docs (I didn't say they should create docs but I think the other models looked at previous commits and figured out they should do that?) - Ultra was the only model that didn't prominently display the "give me the grant to upload assets at the end" - I think Opus 5.5 is the best result. It really feels like more creative (headline) and it nailed the screenshots. - High is my second favorite.
23
2
118
12,963
I'll say this: I've gotten more out of using orbs and asking agents "test this e2e & give me irrefutable proof it works" than out of 80% of tests I've written and ran over the last 15 years. It's truly remarkable that we can now ask machines: "Did you manually test it?"
I agree with @thorstenball, I think unit tests are dead in the water. The ones the models write are terrible, at best just doubling total LOC. Inverting the testing approach - heavy e2e/black box/golden master, reaching for lower levels only if necessary, works better for me
40
8
357
30,717
"this technology will create little moments of delight, whimsy, and value, in everything we touch" yes!
people don’t seem to get the idea that this technology will create little moments of delight, whimsy, and value, in everything we touch. the whole reason you don’t fine tune a classifier for one task is so users can shape the software to themselves, not vice versa
1
2
47
7,477