Co-founder @Calibre_Labs | Applied AI research & consulting | Agents, AI Evals | prev EVP @amplitude_HQ | VC @khoslaventures @sequoia | @stanford @iitbombay

super helpful! i have literally never used low effort for obv reasons and will try it out today for storming a new idea, was definitely missing this interaction type with how eager 5.5 is to get shot done
Replying to @trq212
I use effort low a lot more when I want to stay in the loop, and effort max is basically only if I want zero input or want to find security vulnerabilities.
1
416
Loved the recent article on evals from @HamelHusain and @sh_reya in the @lennysan newsletter. One question we get everytime with PMs is what evals to write for long horizon agents that do a broad swath of work (what H/S call error discovery, similar to product discovery). Below is a great framework for how you can organize these discoveries as a product team: 1️⃣ Outcome - were the results great 2️⃣ Trajectory - was the process followed correct 3️⃣ Experience - did it feel right 4️⃣ Governance - did it follow critical rules You start with your product pov, but the only way to hone focus is to look at traces (trace analysis). Here's an example of an annotated trace for a local shopping assistant called Corner and how often the trace reveals flaws before they become obvious to the customer. claude.ai/artifact/PBvMFLueC…
7
15
46
3,872
Full articles: Eval Rubric: blog.calibrelabs.ai/p/an-eva… Advanced Evals preview: lennysnewsletter.com/p/advan…
1
1
185
I don't know how much of it is a "no mannered prose" instr and how much the post-training but coding with Opus 5.5 feels spectacular in terms of comms. It's also a very eager model, jumping on tasks before I'm quite ready to act on my plans, getting used to life on the edge :/ I've been pushing it on product taste and have noticed giving it a small sample of your strongest opinions works wonders, it comes up with better ideas than most PMs tbh.
1
3
459
Can we agree to call Jev a "Decision Model"? I feel like they buried this in the explainations and docs. Language models output language. Decision models output decisions.
48
12
302
79,852
ML is coming full circle with specialist cheap fast models that will be tiny and live in browsers(?) - Jev is going to be a big deal, get on it if you haven't tried it yet! I am loving it for the evals use case. You know I don't hype random influencer launches, only the good shit. :)
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
3
1
12
2,486
The bitter lesson of AI UI design
Claude Cowork and chat are merging into one Claude. Ask a quick question or hand over a report, and Claude takes it from there, even after you close your laptop. If something's unclear, Claude asks—you keep the final say. Rolling out to Pro and Max over the next few weeks.
1
1
4
669
easily #1 issue amongst teams we work with - no one is reading.. incredible amounts of alpha in being detail oriented rn!
2
3
775
summoning @aptshadow
the fly brain can play beat saber
389
been a while since i listened to an entire podcast episode and really enjoyed both Walter Russell Mead and Greg Jensen as podverse guests this week .. still about AI but from a non silicon valley perspective and provoked thought about its impact on geopolitics and finance respectively.. Look them up!
1
226
idk why but folding phones have an unc vibe i don’t make the rules
6
1
12
1,671
About two years after the launch of o1/reasoning models.. 6 months after ARC-AGI-3 itself was launched we have a 100% score FC finds legit.. definitely a big moment.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: arcprize.org/blog/astra
1
3
686
Loved working with Amplitude on this killer product! It makes no sense to have an online evals stack that doesn't include product analytics data - end of the day, user engagement and retention is one of the most important reward signals for AI agent experiences. This also bring agent experience closer to PMs, designers and domain experts. When builders who deeply know their customers can contribute to AI development in companies, the results are so much more compelling.
Today we're launching Agent Analytics Every team shipping an AI agent has the same blind spot. Offline evals pass, you ship, and then you have no idea what's happening in production. AI fails silently. Users ask a question and get different answers. They all look 'engaged' in a classic dashboard. You don't know who got a great response and who got a terrible one. Agent Analytics solves it: - Every session scored out of the box on task completion, response quality, friction, safety, and negative feedback - Topic clustering across thousands of conversations, so you know if a failure hits 1 user or 10,000 - Eval agents that watch for regressions, and if you want will file a Linear ticket or the pull request themselves - Agent quality sits next to product data, so 'payment scheduling fails 31%' becomes 'which renewals did that cost us?' The Economist got their agent to a 96.9% task success rate and cut weekly failures 84%. Included on every plan. Free tier included. amplitude.com/agent-analytic…
2
1
8
877
New users of OpenClaw in the US fell from an est. 730k in March to <50k this month. It’s still a very important project but a reminder that you don’t need to treat each launch like you are going to be left behind by the whole world. If a new buzzy product is largely being promoted by its investors and influencers, bide your time. They have an incentive to spend their time kicking its tires in public - you absolutely don’t. Your FOMO and anxiety is their engagement. Tuning out social media noise and learning how to build good AI workflows yourself on a platform you already use amounts to a basic survival skill now.
2
2
24
2,294
Hanging out with the original Claude today
5
342
Woke up with PHASEONE[big] energy today need to orchestrate a heist
1
4
470