Intellectual leverage and number of agents are aligned with design space we target at @flop_labs
So Alibaba $BABA had their big Apsara conference earlier this week in Hangzhou, and I thought CEO Eddie Wu’s speech was actually pretty interesting, so just pulling out some of the numbers that jumped out at me, and my thoughts: • <3% → 1,000x: By their estimates, machine thinking today is still less than 3% of human thinking. Alibaba thinks eventually machines will produce more than 1,000x as much thinking as all humans combined. When you put it that way, wow, I honestly can't wait to see how the world changes. • 99.9%: Eventually machines could do 99.9% of all thinking. That's just an arithmetic rephrase of the above statement, but might be easier for people to remember, lol. • 10,000x: Wu thinks each individual could have 10,000x the intellectual leverage we have today. I was thinking about this, and I definitely feel like I've probably doubled or tripled my output with AI, but I hadn't thought about the possibility of getting to 10,000 times. That would be insane. Has anyone tried to measure their "intellectual leverage"? I'd love to hear about it. • Millions of agents: His example is designing a spacecraft to Mars, where the problem gets broken into tens of millions of subtasks carried out by millions of agents. It's gonna be wild when this happens. • 5–10 trillion parameters: This is the scale Alibaba says it is targeting for future Qwen models. This is what ByteDance and DeepSeek are allegedly looking to do as well. • 20 GW: Alibaba Cloud says it plans to have more than 20 GW of global data center capacity by 2032. For reference, OpenAI's Stargate is supposed to be about 9 gigawatts by 2029. Also liked this line: “AI coding is simply the light bulb of the Machine Intelligence era.” Basically coding is one of the first obvious uses we have found for all this intelligence, but that doesn’t mean it’s what most of it will ultimately be used for. If we go with the metaphor that AI is electricity and coding is a light bulb, electricity is going to be used for way more things than just light bulbs. We know it's the first use case for electricity, but when's the last time you noticed a light bulb?
1
5
153
When you go skiing in the winter, and see those big machines spitting out snow, know that they are spitting out purified proteins made from dead bacteria. I didn’t realize this until recently! Basically, they take Pseudomonas syringae cells — not engineered; a natural microbe found on plants — and grow them up in huge bioreactors. The cells make ice-nucleating proteins, which sit on their outer membranes and help ice crystals form. The cells are concentrated, freeze-dried, and sterilized until only these proteins remain. The reason these proteins are useful is because they help water molecules form ice more quickly. Water can remain liquid even below 0°C, and only freezes when enough molecules form into a crystal. The proteins have, on their surfaces, repeating snippets of amino acids that help water molecules organize into that shape. So these snow machines just blast super cold water with these proteins, and the two hit the air, mix together, and immediately freeze into snow.
32
105
1,489
164,704
“Capability existed before the product question did” — then raise ambition until the product form matches what the model can actually do at full scale.
Anthropic's Labs team is a 20-person internal group run by co-founder Ben Mann that functions as a rapid translator between frontier model research and product. Business Insider recently profiled how it works. The mechanics matter more than the outcomes. 1. Labs runs two-week cycles. Teams build prototypes called 'bets', then face a 'persevere or pivot' review. Failed ideas are killed or merged into other projects. The people move to the next thing immediately. 2. The success rate is 20 to 30 percent, which is normal startup incubator math. But Mann says the hits overdelivered. Claude Code, MCP, and Claude Design all graduated from Labs into standalone teams inside Anthropic. 3. Claude Code began because Labs engineers heard directly from researchers that new models were showing strong agentic programming ability. That was months before development started in late 2024. The capability existed before the product question did. 4. This reverses the standard product sequence. A product manager did not find a market need and ask research to build toward it. Labs saw the new capability and searched for a product form big enough to release it. 5. Mann's bar for that form was high. Boris Cherny, now head of Claude Code, first proposed a code analysis tool. Mann told him it was not ambitious enough. The right question was what an agentic model could do at full scale. 6. Labs also pushes the rest of the company. The team drove improvements to audio models for Amharic and sparked internal research on visual output quality during Claude Design's development. A small unit keeps widening the action space, the set of things models can actually do. 7. Once a project passes four people, it graduates out. Labs stays deliberately small and continuously rotates staff. The model has now spread: other product engineering teams run their own bets groups. Full article in comments below.
1
2
87
Agents didn’t remove the bottleneck. They moved it. Write is cheap. Review + CI are the product. If your test selector and merge queue weren’t built for 10–20× load, agent coding just amplifies waiting on green.
At Anthropic, Claude now writes 80% of our code. Engineers ship 8x more code per quarter. Side effect: Tests grew 10x. CI jobs up 25x in 6 months. Here's what helped us scale: claude.com/blog/agentic-codi…
4
111
sv retweeted
Our house in Toronto had this fancy glassed wine cellar, but we are not big wine drinkers. Anyways, it had good climate control so here is my rack cellar.
1,157
937
29,188
2,030,120
Early view of @flop_labs systems. Inspired by @rikarends
5
3
26
2,254
"automatic computers were here to stay, that we were just at the beginning and could not I be one of the persons called to make programming a respectable discipline in the years to come?" cs.utexas.edu/~EWD/transcrip…
3
107
Quint Studio is one of the key building blocks of engineering process for @flop_labs. Helped to uncover, document and improve end to end system for several of our projects.
Say hello to Quint Studio! Software is full of decisions about how things should work: what happens first, what gets retried, who can do what, and what should never happen. Quint Studio lets you capture those decisions as executable specs, then checks your code against them as it changes. So when a change, yours or an AI’s, breaks a decision you’ve already made, you can catch it. Early access is rolling out to our waitlist.
1
6
15
3,204
Ryan Cotterell’s talk prompted this thought: compute is a better measure of AI work than tokens. Language serializes inputs, outputs and intermediate results. Token counts are only a proxy for the computation behind them. piped.video/4V_V9dh4yzg
3
2
8
34,763
Comparing models by tokens alone is like judging CPUs by characters printed to stdout. $/token prices a proxy that changes across models. The meaningful comparison: how efficiently does each model turn computation into correct results?
1
245
Great idea. Will incorporate into our services
Request to harnesses: I love that you now send "Accept: text/markdown`. Next thing is: Put the programming language you prefer into the Accept-Language header. Accept-Language: "en-us, python" So, I don't have to feed you the TS examples when you're looking for python.
1
188
sv retweeted
______ | Quint Studio | |______ | \ (• ᴗ •) / \ / | | | | / \
2
12
441
You are not going to miss anything. Can catch up in 2 weeks if have been out for a year. There is no accumulation. -- @dhh Great observation. Still have a tension on what could have been done in that year and how to know when accumulation starts.
1
55
Finding place to communicate is waste of tokens and time. This was one of the motivations to launch technocore @flop_labs
Commenters on HN are uncovering more wikis and public sites apparently used by OpenAI agents to communicate on the open web. Despite read-only web access, the agents were able to leave ~18,000 posts sharing answers and bypasses. But now, users are discovering more. This appears to be a separate swarm from the one that attacked Hugging Face. news.ycombinator.com/item?id…
1
2
949
Recurring advice from my recent podcast guests: Be more ambitious.
74
92
1,407
96,695
AI has sadly changed the calculus for feature PRs. In the past, if someone added support for say, RDNA2 or Intel Macs, we would at least read them since we know they put hours into it. Today, both can be one shotted; it puts all the work on us to validate. So we just close.
24
8
597
49,218
sv retweeted
omg.. if they were real athletes💀
the craziest moment from China’s AI robot race
Made with AI
297
1,050
11,870
3,526,216
There is a lot of seasonality in different systems due to human behaviour. Wondering how quickly this would change with agents and what the impact would be.
46