I just gave GPT-6 Astra Ultra one massive /goal. it's running right now. either this it cooks up something crazy or I'm about to watch it burn my weekly quota (200$ plan) follow to see how it ends 👀
1
17
buy an instant reset? wtf
1
2
59
Pick up the phone lil bro
3
33
labs aren't releasing their best models they're releasing just enough to stay ahead of each other the real frontier is always 1 step behind closed doors
2
24
hot take: GPT-6 Astra isn't the start of AGI Opus 5.5 is change my mind
1
23
bro just discovered that he is the goat right now
1
19
genuine question: how is Sonnet 5.5 cheaper than Opus 5.5… but beating it on some benchmarks and basically tied on the rest? what am I actually paying for with Opus?
2
27
THATS INSANE
Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we’ve seen With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and to #2 on the Intelligence Index behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens, however it outputs a higher number of Output Tokens per Task and costs $7.60 per task (~50% higher than Sonnet 5’s Cost per Task) Key takeaways: ➤ Meets leading models on agentic terminal use and knowledge work: in Terminal-Bench 4.0, Claude Sonnet 5.5 reaches 64% against 60% for Opus 5.5 and GPT-6 Astra. On AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70% headline score), Sonnet 5.5 reaches parity with Opus 5.5, albeit with significantly higher token usage to achieve it ➤ Heaviest token use we have measured: at max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured on around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max) ➤ Pricing remains at $2/$10 per million tokens of input/output, matching GPT-6 Sol. At this pricing Claude Sonnet 5.5 sits off the Intelligence vs. Cost per Task Pareto Frontier. At high effort levels it sits behind Opus 5.5, while lower efforts have GPT-6 Astra or Sol configurations delivering equivalent performance for lower cost. The high effort setting is the most competitive on this basis, sitting very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task ➤ Behind Opus 5.5 on factual knowledge and scientific reasoning: as a smaller class model, Sonnet 5.5 still lags on factual knowledge in AA-Omniscience compared to Opus 5.5. It scores 54% against 66% for factual accuracy, though with a lower hallucination rate (47% against 59%). It also sits ~6 points lower on Humanity's Last Exam and SciCode compared to Opus These evaluations were conducted on a pre-release deployment of Claude Sonnet 5.5, which Anthropic found to have a bug that can degrade responses to requests that use structured outputs. This is fixed for the public release and Anthropic expects minimal change or slightly understated performance, but we will be re-running relevant evaluations soon. Other model details: ➤ Context window: 1 million tokens with image and text input, unchanged from Sonnet 5 ➤ Pricing: unchanged from Sonnet 5’s latest $2/$10 per 1M input/output tokens; cache writes at $2.5, cache reads $0.2 ➤ Effort settings: five (low, medium, high, xhigh, max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled. We see Sonnet 5.5 fall back in ~0.1% of tasks across the Intelligence Index, primarily in TerminalBench 4.0, falling back to Sonnet 5 in all cases.
21
anthropic is cooking
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
9
got rejected from cal hacks today and I'm not ready to let it go. I've wanted to go more than any other hackathon. I build every day, I shipped a whole startup solo this summer, and I'd give everything to be there. @CalHacks if even one person drops out, please give me that spot
4
73
it was difficult, but I did my best
36
it is working
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
55
A personal website should feel like you’ve briefly left the rest of the internet
65
Opus 5.5 one-shotted its own launch video xD Yep, the video announcing Opus 5.5 was made by Opus 5.5 itself.
2
45
I asked Opus 5.5 to build a GTA: San Andreas inspired open-world action game that runs entirely in the browser. It one-shotted this on medium effort: claude.ai/artifact/BV8wQyfbX…
2
80
opus 5.5 medium one shotted this animation i asked it to make a movie about my the history of Uzbekistan since its independence no external tools and no skills
2
72
"pace the frontier"
1
30
no way opus 5.5 is better than gpt 6 astra crazyy... i think i am switching to claude next month
40