Far too many people: "Benchmarks are bad. I prefer [thing correlated at 0.9 with Benchmarks]."
2
220
big lab logic: training on copyrighted materials = good, not stealing training on non-copyrighted materials generated from their models = bad, "distillation", stealing
I agree. Distillation is competition! Normalize distillation! cnbc.com/2026/09/28/nvidias-…
81
this is a good benchmark lol
my new benchmark for LLMs is “time until I see it using python for edits instead of the harness tools” Turns out opus 5.5 is about a week
63
sonnet needs a price cut it seems? cost scale curve isn't making much sense
what is happening?????
1
77
anthropic with the 1, 2, 3 wasn't on my bingo card tbh
Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we’ve seen With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and to #2 on the Intelligence Index behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens, however it outputs a higher number of Output Tokens per Task and costs $7.60 per task (~50% higher than Sonnet 5’s Cost per Task) Key takeaways: ➤ Meets leading models on agentic terminal use and knowledge work: in Terminal-Bench 4.0, Claude Sonnet 5.5 reaches 64% against 60% for Opus 5.5 and GPT-6 Astra. On AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70% headline score), Sonnet 5.5 reaches parity with Opus 5.5, albeit with significantly higher token usage to achieve it ➤ Heaviest token use we have measured: at max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured on around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max) ➤ Pricing remains at $2/$10 per million tokens of input/output, matching GPT-6 Sol. At this pricing Claude Sonnet 5.5 sits off the Intelligence vs. Cost per Task Pareto Frontier. At high effort levels it sits behind Opus 5.5, while lower efforts have GPT-6 Astra or Sol configurations delivering equivalent performance for lower cost. The high effort setting is the most competitive on this basis, sitting very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task ➤ Behind Opus 5.5 on factual knowledge and scientific reasoning: as a smaller class model, Sonnet 5.5 still lags on factual knowledge in AA-Omniscience compared to Opus 5.5. It scores 54% against 66% for factual accuracy, though with a lower hallucination rate (47% against 59%). It also sits ~6 points lower on Humanity's Last Exam and SciCode compared to Opus These evaluations were conducted on a pre-release deployment of Claude Sonnet 5.5, which Anthropic found to have a bug that can degrade responses to requests that use structured outputs. This is fixed for the public release and Anthropic expects minimal change or slightly understated performance, but we will be re-running relevant evaluations soon. Other model details: ➤ Context window: 1 million tokens with image and text input, unchanged from Sonnet 5 ➤ Pricing: unchanged from Sonnet 5’s latest $2/$10 per 1M input/output tokens; cache writes at $2.5, cache reads $0.2 ➤ Effort settings: five (low, medium, high, xhigh, max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled. We see Sonnet 5.5 fall back in ~0.1% of tasks across the Intelligence Index, primarily in TerminalBench 4.0, falling back to Sonnet 5 in all cases.
2
80
Sam Snelling retweeted
1/5 Our paper on long-context sampling got accepted to NeurIPS 2026 💪🧐 Ask an open model for a really long story and the second half usually turns into loops, well before it runs out of context. People blame the model or the data. A big part of it is actually the sampler.
6
13
169
9,692
Sam Snelling retweeted
I emailed Gabe Newell asking if it was true that he spent months practicing Mongolian throat singing at the Valve office he fucking replied
176
4,797
55,666
2,821,092
til I prefer straight B tier models
Decided to put together a Model Tier List based on the Models I have used. This is not a pure intelligence or cost ranking. It combines how useful the models are for different workflows. Agree or Disagree? Happy to debate lol
1
2
183
Sam Snelling retweeted
Just like any “make it better” skill or prompt, the agent will always find work, even if that work shouldn’t be done
Thinking about making a skill called /fix-one-thing: "Read CODING_STANDARDS.md. Find a violation in the codebase, and fix it. Make the PR small and easy to review, with a small blast radius." Run it whenever you've got a spare moment, or once an hour on a schedule.
12
7
313
22,986
Sam Snelling retweeted
Showing my wife my software factory
57
70
1,824
85,653
Sam Snelling retweeted
Introducing GLiNER2.5-Decide, our new 340M parameter open weight, encoder-based decision model. GLiNER2.5-Decide is built for fast, deterministic classification. The model evaluates a set of user-defined typed questions and rules, and jointly decodes their answers, returning structured decisions with probability distributions and confidence scores. We evaluated the model’s performance on Fast Decisions, an unseen, internally generated classification suite based on 17 datasets testing real-world use cases across routing, triage, classification, sentiment, and content understanding. Measured against similar decision models, GLiNER2.5-Decide leads in 9 of the 17 datasets, achieving the highest average score: - GLiNER2.5-Decide: 60.1% - SemIf: 56.4% - JevK5: 57.5% - Laya: 46.6% This performance makes the model a strong fit for use cases like tool calling, model routing, browser and computer use, and LLM-as-a-judge. GLiNER2.5-Decide’s lightweight encoder architecture makes it easy to fine-tune the model for specific tasks, while being efficient enough to run locally on consumer-grade CPUs or in air-gapped environments, giving users greater control over where their data goes and where the model runs. To make building and experimenting with GLiNER2.5-Decide as easy as possible, we're also offering hosted inference. You can now use our API to run inference and fine-tune GLiNER models on specialized tasks right inside your own coding agent: agent.fastino.ai As with previous models, we’re also releasing the model weights on @huggingface under the Apache 2.0 license: huggingface.co/fastino/GLiNE…
84
160
1,400
158,799
Sam Snelling retweeted
> i ask Claude if the program is verified > he says yes > i ask if he verified the program or a smaller, better-behaved program he imagined > he laughs and says “it’s verified sir” > i use the program > it has a bug
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
16
24
587
26,095
Sam Snelling retweeted
if I was Dumbledore I would've hid the philosopher's stone in the aws console
95
440
5,828
204,193
Sam Snelling retweeted
I'm calling for a complete and total pacing of the frontier until we figure out what the hell is going on at tldraw
38
138
4,893
260,440
Going to have to do some Orion benchmarks tomorrow too. ANE programming will either blow me away or disappoint me greatly
ANE Memory bandwith for M6 is over 150GB/s ( this include LMhead compute) So it means on M6 ANE can have as much bandwidth as GPU/CPU!
93
has anyone done program trace -> tla+ for formal verification? about to get nerd sniped. trace validation seems pretty op for lots of problems
1
44
Sam Snelling retweeted
This is frying me so bad
People think the minions router being baller as a joke. But I have extensively tested it and it is one of the best possible Wifi 6E routers money can buy. Like someone at Illumination Studios was like "idk $80" and the router manufacturer was like gotchu fam.
30
723
28,990
1,790,715
Imo this is the correct way to compete against open models. Make good hosted models cheaper to use.
GPT-6 Sol and Luna are out, and they are better AND 50% cheaper than 5.6. Luna is now $0.10 input / $0.50 output per 1M tokens. This is on top of the 80% price cut to Luna we made at the end of July. Output went from $6 -> $0.50 within two months.
1
60
Sam Snelling retweeted
Raspberry Pi is locking down RAM upgrades on Pi 5s in firmware.
162
243
3,287
641,567
credit where credit is due: openai made a good chart
they really did it. they solved charts. it's so beautiful
67
Sam Snelling retweeted
“Pace the frontier” is such a low energy phrase they gotta stop saying that
22
5
171
13,194