new model - please enjoy
13
2
186
13,504
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
4
3,195
wat
What the actual fuck. Opus 5.5 ultra created this masterpiece. 4 agents and an hour and a half later. Everything from scratch, no AI voice API used... this is it...
1
32
12,707
Adam Feldman retweeted
Claude Opus 5.5 drew every frame of this animation in JavaScript. Everyone in town sends Claude their requests, but one girl sends a question instead: "What do you love?"
227
382
6,077
644,019
Adam Feldman retweeted
At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent. Overall, we saw frontier capability levels, with major performance improvements over Opus 5. 63% fewer tokens used, 42% less verbosity, and 30% faster vs. Opus 5. And the model itself is cheaper, so this is a major win for any agentic computer use, coding, analytics, or data work that enterprises will be doing. Here are some examples of the task wins and performance gains across a variety of industry tests that we performed: • Financial services - due diligence (+39% task accuracy): A year of transaction records, with the job of finding every miscalculation in an acquisition target's pricing tool. Opus 5.5 scored a perfect result on every attempt in half the words Opus 5 used, consuming 82% fewer tokens overall. • Technology - cloud cost analysis (+65% task accuracy): Work out what a company should actually change about its cloud spend. Opus 5.5 picked the right basis for the retention calculation and kept the source data's unit conventions straight all the way through, so the number at the end actually holds up. It took half the time Opus 5 took, with 70% fewer tokens. • Consumer products - client account analysis (+17% task accuracy): Set the onboarding targets for a client account, reading across the signed contract, a satisfaction tracker and a team metrics sheet. The contract never states a senior/junior split, so Opus 5.5 derived it from the 18-person roster and showed the rule it used; several clients had a perfect 10 on individual survey questions, so it averaged each client's responses instead of crowning the single 10. It finished this one in half the time, on 78% fewer tokens. • Clinical diagnostics - data analysis (+15% task accuracy): Malaria rapid-test performance across a dry and a wet season: build the patient records out of two clinical PDFs, compute positive test rates by season and gender, and test whether parasite counts really differ between test-positive and test-negative patients. Opus 5.5 caught that the two groups' standard deviations differed more than 100-fold, re-ran it the right way, and found the dry-season difference didn't hold up after all. This accuracy gain came with a final answer that was half the length of Opus 5's, and also needed 78% fewer tokens end to end. Customers will be able to build AI Agents with Opus 5.5 shortly in the Box AI Studio.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
60
28
388
152,685
Adam Feldman retweeted
also important news we fixed the writing
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
174
114
4,237
345,939
"Americans still view themselves as the victims of 9/11."
2
6
138
9,878
We combined the things. Please enjoy.
Claude Cowork and chat are merging into one Claude. Ask a quick question or hand over a report, and Claude takes it from there, even after you close your laptop. If something's unclear, Claude asks—you keep the final say. Rolling out to Pro and Max over the next few weeks.
9
64
5,926
Army Now Testing Anthem Long-Range Strike Weapon Built for Mass Production Serial production of the relatively crude but cheap weapon is slated to begin in Texas next year, with plans to also build the weapon in Germany and Israel. twz.com/news-features/army-n…
1
7
1,336
launched another model - please enjoy
10
99
5,382
Adam Feldman retweeted
really feel this sometimes
59
462
11,530
531,676
Adam Feldman retweeted
Rep. Dan Goldman (D-NY), in concession speech: “Jews have given back so much to this country. As history has taught us, antisemitic tropes and stereotypes, some of which I heard personally on this campaign, will ultimately be the undoing of our democracy if we all don’t lean in and speak out — even if it’s not politically expedient.”
625
728
4,951
855,783
Adam Feldman retweeted
This is a new paradigm for interacting with Claude that is significantly more "inline" with all the other human activity org-wide. Once you do all of the under the hood engineering work to make this "just work" (e.g. across tools, integrations, compute environments, memory, security, etc.), Claude basically joins the team in a seamless way - you can talk to it as you would talk to a person and it can help with a very large variety of workloads. Imo this is the 3rd major redesign of LLM UIUX. The first paradigm was that the LLM is a website you go to, the second was that it is an app you download to your computer. This third one is that it is a self-contained, persistent, asynchronous entity with org-wide tools and context, working alongside teams of humans. It really takes a while to wrap your head around it, but it works and it is awesome.
Introducing Claude Tag, a new way for teams to work with Claude. In Slack, Claude joins as a team member with access to the channels and tools you choose. Tag Claude in and delegate tasks to it while you focus on other work.
1,381
1,941
23,464
8,369,079
Adam Feldman retweeted
Our Labs team worked on this along with the Claude Code team — as Claude's work gets deeper and more independent, I've found really valuable to have it explain its thinking and outputs using Artifacts; try asking it to diagram its work next time you're working with it.
New in Claude Code: Artifacts. Interactive pages built from your session, like a PR walkthrough or a living project dashboard, shared with your team at a private link. Available in beta on Team and Enterprise plans.
15
6
213
27,904
Adam Feldman retweeted
Our most capable model is now available everywhere, including our mobile apps. Let us know your thoughts!
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available.
15
3
138
8,125
Adam Feldman retweeted
My favorite chart from our system card - FrontierCode is an excellent eval, and it accurately reflects the step up I feel when using Fable!
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available.
35
39
626
66,550
We just released Claude Fable 5, a Mythos class model made safe for general use. This is a big one. I'm thrilled that everyone can now use the capabilities we've had internally at Anthropic these past few months.
9
87
3,613
Let us know what you think.
hello beloved tasteful users, do you like how much claude thinks on your tasks? would love examples of it thinking too much or too little
1
4
1,130