making computers do more with Call It

why?
my agent can now handle full multi-step tasks in the background in this demo, it researches products, compares them in Notes, picks the best one, adds it to Amazon, and drafts an email with the link all while I keep using my computer normally
6
10
32
10,891
another day another claude
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
5
129
might just start using Luna for a bunch of things i’m building now. $0.10 in / $0.50 out per 1M is pretty cheap
1
1
5
164
Opus 5.5 is ahead of GPT-6 Astra on the AA Intelligence Index and costs 2.5x less that’s a pretty serious gap
4
172
Opus 5.5, GPT-6 Sol, GPT-6 Luna all in a single day nobody is slowing down
1
4
139
fable 5.1 performance for 40% less than opus 5 is kinda crazy!!
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
3
140
I'm attending Syndicate, @aoagents' online hackathon. Cash prizes are up for grabs, powered by @maximor_ai and supported by @dodopayments, @neatlogs, @Tensormux, and @aigrantsindia. See you there! aoagents.dev/hackathons/synd…
1
7
132
Fable all day now!
BREAKING: Anthropic just dropped Fable 5.1—and CLAUDE IS SO BACK. We’ve spent the last week testing it at @every across coding, writing, and knowledge work. Our verdict: It's finally Fable for everyone. It’s the strongest coding model we’ve used, but now it's fast, token-efficient, and CRUCIALLY actually speaks like a normal person. Here’s our vibe check: - A monster at coding. @kieranklaassen rebuilt a working version of Proof, our document editor, from one prompt. It added useful details he hadn’t requested, and it handles enormous coding jobs that run for days at a time. It built a computer use Mac app for me called Hands in one-shot that other models failed at. - A Claude our writers want to use again. It has clearer prose, fewer AI tells, and it takes an edit without arguing. It's a significant upgrade over Opus 5. And won @kplikethebird's heart back. - About half the tokens as Opus 5, and much faster. In our Slack-agent tests, it delivered comparable results to Opus 5 using about half as many tokens, in about 60% of the time. - Knowledge work you can delegate. It can produce great knowledge work—like slide decks—end to end without making slop. And flew threw @hammermt's tests with flying colors. - It now supports zero-data-retention agreements. Now businesses can actually use it! A big barrier to Fable adoption is gone. Net Result: It's obviously an Opus 5 killer. If that was your daily driver you should switch today. If you're using GPT-5.6 in ChatGPT for Work, it's spinning the wheels on for big delegated tasks. I still use ChatGPT for Work more day to day, but I use way more tokens in Fable 5.1. I send it off at the beginning of the day to do big programming projects, like end to end MVP builds, and check in every once in a while. State of Play: The big knock on Anthropic was they built a supergenius in a datacenter that was almost unusable. It was too slow, argued back, and talked in technical gibberish. They've managed to solve those problems and more with Fable 5.1!
1
2
272
more brains, less burn. if this actually holds up in real use.. they cooked!
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
1
121
learned about RoPE today and why it’s a better way to handle position than plain positional embeddings normal positional embeddings basically tell the model where a token sits in the sequence RoPE handles that inside attention itself by rotating Q and K based on position so when two tokens are compared, the model gets not just the content match but also how far apart they are and which one comes before the other pretty cool how position ends up becoming part of the attention itself instead of just being added on top youtube.com/watch?v=hCzJo4ui…
7
8
118
got a bit into attention today Q and K are doing the matching, basically figuring out how much each V should matter take that weighted mix of the Vs add it back to the token's vector and now it has context from the tokens that mattered to it so yeah pretty much tokens getting contextualized about their surroundings
7
8
113
Divyendra Singh retweeted
ig this is how we contribute now
We’ve decided to open-source a multi-agent harness we use internally at YC. We call it “QM” and it’s meant to be easy to customize, like Hermes or OpenClaw, but useful for a whole company. We use it across accounting, legal, events, and engineering (including building QM itself!). The whole project is under an MIT license. It is cloud-first and has Slack and web UI natively.
9
11
348
Divyendra Singh retweeted
made my desktop agent work as an MCP for Claude Code now it can pass exactly what I highlighted, where I pointed, what’s on screen, and what I said all directly into the coding session no more taking screenshots, cropping them, and explaining everything again
2
10
16
1,452
Divyendra Singh retweeted
made my agent much faster at repetitive tasks instead of calling a large model and figuring everything out again, it can recognize workflows locally and run them directly I initially thought about letting users demonstrate a task and turning that into a skill, but the way a human does something is rarely the best way for an agent to do it less reasoning, fewer round trips, faster execution
3
7
11
2,071
nice to see @every doing these vibe checks right after launch really helps figure out how a new model actually behaves
BREAKING: Claude Opus 5 is OUT NOW! And…it’s a hard model to love. We’ve spent the last week @every testing it across coding, writing, knowledge work, and our internal agent. It argued with instructions, stopped before the work was finished, and generally didn’t play well with our existing skills and plugins like Compound Engineering. Our first reaction was: What have they done to my boy? Then we deleted our existing skills and started from scratch. Without the elaborate workflows we had built for earlier models, Opus 5 got dramatically better, and even showed flashes of brilliance. Here’s our Day 0 vibe check: - It’s a poor man’s Fable. It has many of Fable’s personality quirks without Fable’s genius. - It breaks backward compatibility. If you’re using it with existing skills and workflows, watch out. It will often stop early or otherwise miss your instructions. - If you start from scratch, you’ll have better results. @KieranKlaassen figured out that if he just started from scratch without his existing skills, he could get dramatically better results. This is a model that takes some time to rebuild your workflows around—but if you do, there’s a payoff waiting. - Medium or low effort works better. @KieranKlassenn also found better results using Opus 5 on lower thinking levels. It seems that the more time you give it to think, the more likely it is to do the more annoying behaviors. Don’t just switch to Sonnet for a faster response! Try low thinking. I have two slots in my workflow: 1. The genius model I use for my biggest hardest tasks, currently Fable. 2. The smart, fast generalist I use for everything else, currently GPT-5.6. Opus 5 has the personality of the genius, but doesn’t have its top end. So that puts it in a strange middle ground that doesn’t really have a home in my day to day. I think I’ll use it mostly when I run out of Fable tokens. full vibe check on @every in the next tweet 👇
3
5
394
close to fable at half the price is probably the only benchmark most people needed
Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.
5
8
252
the weird part about building agents today is that they can be surprisingly capable, but they still don’t really learn from using your computer you have to build that layer yourself
Humans try hard things, fail, learn, and get better through repetition. AI models aren't as different as you might think. Here's how models "learn" explained in simple terms.
Article

How we teach AI models

How do AI models learn new skills and behaviors? The process is surprisingly human and easy to understand, even if you don’t have a machine learning background. Helpful colleagues AI models,

6
9
156
the part about smaller prompts not always being cheaper was interesting. aggressively pruning agent history can cost more than just keeping the cached context around
We've recently made Pi's cache behavior more visible. This site has been debating whether agent harnesses are helping or quietly torching their caches. That seemed like a good excuse to explain how KV caches actually work and how Pi helps (or doesn't). earendil.com/posts/prompt-ca…
2
4
198
yeah, I’ve been doing the same. voice helps you give way more context than you’d ever bother typing. it might come out messy, but the model usually picks up the important details
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
2
3
233