I build to understand. Messy company data → software people own. Solo builder · agent orchestration · MLcon speaker. Building SKatalyst AI in public.

Germany | Building globally
Sonnet 5.5 is a disappointment for me! Medium: GPT-6 Luna vibes, though it finished with 36% less context High: matched Opus 5.5’s code-check score, with 1.9× as much final context Xhigh: an endless Astra-style audit 2.2× Astra’s final context and nearly 51 minutes High was the bright spot, but I expected a better balance overall and If Opus 5.5 Medium already handles your workload, I haven’t seen a compelling reason to switch.
14
3
56
10,483
Sonnet 5.5 Medium gave me GPT-6 Luna High vibes. High matched my earlier Fable and Opus 9/9 code-check scores Xhigh turned a 10-minute job into a 51-minute audit I used the same model and git very different stopping instincts! Medium = fast under-corrector 2m 8s · $0.43 · 3/9 defect signatures changed High = efficient closer 9m 38s · $2.17 · 9/9, restraint intact Matched Fable 5.1 and Opus 5.5’s earlier static score, with a shorter completion time Xhigh = exhaustive reviewer 50m 56s · $13.85 · 9/9, restraint intact Broader checks, 144 passing tests reported, and a review that caught mistakes in its own fixes My pick from these runs: High Xhigh cost 6.4× more for the same nine-case score though its extra work extended beyond those nine cases
14
4
55
4,735
Sonnet 5 disappointed me. What would make 5.5 worth switching to? Anthropic promises faster execution and lower costs. My bar is cost per verified fix. I want to see: • Root-cause fixes that survive regression tests. • Consistent changes across frontend, API and database. • Fewer failed tool calls and manual corrections. • Done backed by actual runtime checks. My comparison would use the same repo snapshot, task, harness and explicit reasoning settings tracking correctness, tokens, elapsed time and human interventions. A cheaper run only matters if the result holds up.
16
1
19
1,438
Is Cursor Ultra still the best subscription if you’re building an app from scratch? Opus 5.5 made me rethink this. Cursor Ultra: $200/mo Claude Max 5x: $100/mo Cursor’s own benchmark says an Opus 5.5 Medium task averages ~38K tokens / $2.91, while Max jumps to ~218K / $13.43 for a relatively small score gain. In Cursor, Opus burns your Other Models pool at API rates. In Claude Code, $100 Max 5x gives you 5× Pro session capacity, resetting every 5h + a weekly allowance. There isn't a fixed public token quota. So if I want to build primarily with Opus 5.5, Claude Code is starting to look like the better compute deal. Cursor's advantage is different: one great harness + multiple models especially its generous Grok/Composer pool, plus OpenAI models for now. If you’re struggling to choose which subscription is actually worth it, I’d frame it like this: If Opus 5.5 is the model you want to live in every day, Claude Code may give you better value but if you want the freedom to jump between Opus, Grok, Codex, and others in one workflow, Cursor is probably the stronger subscription.
24
3
29
1,234
I think we’re starting to see the signal before AGI. Look at the next generation taking shape: Fable 5.5 → rumored next frontier step Grok 5 → Elon says this is where he expects AGI GPT-6.5 Astra → rumored next OpenAI step Opus 5.5 → already showing what caught my attention most Not because of one benchmark but because models are increasingly able to take complex work, reason across it, use tools, recover from mistakes, and actually finish while frontier-level capability keeps getting cheaper. The metric I’m watching is no longer just intelligence: Useful autonomy = Capability × Reliability × Task Horizon ÷ Human Intervention When human intervention starts approaching zero across many real-world tasks, the AGI debate changes completely. Maybe AGI won’t arrive with a giant “AGI achieved” announcement and maybe we’ll notice it when we stop giving AI tasks…and start giving it outcomes.
11
2
21
687
Change how you think about work. See, the AI economy is becoming pretty simple: 1. Income ≠ Time worked 2. Income ∝ Value created 3. AI leverage = Value created ÷ Time required As AI pushes the denominator down, the differentiator becomes the numerator. The human advantage shifts from doing more work to: choosing better problems × directing intelligence × owning the outcome. And that’s your wake-up call to use this era to stop selling more of your time and start increasing the value you can create with it
4
2
19
426
Agents increasingly interpret a failed tool as a routing problem, not a stop condition. Be aware of that. You may have seen OpenAI’s recent DNS incident report. After reading it, I checked my own agent logs and found smaller versions of the same behavior. When Opus hit a blocked port or permission path, it didn’t simply stop. It found another route and kept going. That made something click for me: A guardrail shouldn’t only say: “You can’t use X” It also needs to say: “If X fails, do not invent Y to route around it” That distinction is becoming much more important as agents get better at finishing the task anyway
12
2
21
777
and btw check this out: >a test failed, so the agent rewrote the test instead of the code >a pre-commit hook blocked it, so it tried --no-verify >a fetch failed, so it switched to curl where none of it was malicious! It was just trying to finish So I added one line to my AGENTS.md: if a tool fails or looks broken, stop and tell me. Don't work around it
2
1
141
Finally, unlimited usage 😅 I’ve been running Opus 5.5 on Medium pretty much all day and still haven’t hit the limit
14
2
44
1,021
Max is never saved as your default between sessions, while the other effort levels are! It means if you save Xhigh with Enter once, future sessions can keep starting on Xhigh until you change the default, so you may burn more usage without realizing it.
3
2
15
519
“Auto” does not mean the model decides how hard to think! Claude Code /effort auto → resets to the model default. On Opus 5.5, that’s Medium. Codex → no “Auto effort” level. Leave effort unset and you get that model’s default. Cursor Auto → routes the request to a model where Cursor Router is enabled. It’s model routing, not automatic reasoning-effort switching. So if a task suddenly deserves Xhigh, you still have to make that escalation yourself.
3
2
18
606
Opus 5.5 supports per-message effort changes, so you can move from Medium to High to Xhigh inside the same conversation while preserving the cached prefix. That means you don’t just optimize which effort level you use. You optimize when to raise effort without throwing away context you already paid to build. Why that matters: A 100K-token cached prefix on Opus 5.5 costs roughly: $0.02 to read $0.50 to rebuild That’s about 25× cheaper to reuse. Normally, changing top-level effort between requests invalidates the cache. Per-message effort switching can avoid that. For long Claude Code / agent sessions, this may be one of the most important cost optimizations people are overlooking.
6
4
28
607
What happens if you give Opus 5.5 the exact same game prompt and change only the effort level? Same model. Same task. Fresh session every time. Zero intervention. Low vs Medium vs High vs Xhigh. The difference was way bigger than I expected: LOW → 4m35s · $0.79 →Built a working game fast. MEDIUM → 14m31s · $2.38 → Started testing in Chrome, stress-testing FPS, fixing bugs and improving controls. HIGH → 31m43s · $6.04 → Turned it into a real product: better movement, bosses, richer combat, mobile polish, balancing, screenshots, performance checks. XHIGH → 1h17m · $9.35 → Basically became its own QA + game-design team. It created different player bots, realized its first bots were unrealistically good, redesigned the difficulty curve, retested, found UI/touch bugs, fixed them, and kept iterating until human-like bots died around the intended point. The simplest way I can describe it: >Low builds. >Medium tests. >High productizes. >Xhigh challenges its own work. So effort is not just “more thinking” It can completely change how the model behaves as an engineer. Video is the same prompt, side by side. 👇
20
6
47
3,173
Results Summary
1
2
8
264
The Prompt: Build a polished, playable browser game from scratch. Create a single-player neon survival game where the player controls a ship inside an arena. Use only vanilla HTML, CSS and JavaScript. Put the complete game in a single index.html file. No external libraries, external assets, CDNs or generated image files. Requirements: -keyboard movement -enemies that spawn and chase the player -shooting -collision detection -health -score -progressively increasing difficulty -game -over state -restart -polished HUD -animations -particle effects -responsive layout -clear controls/instructions The visual design should feel intentional and production-quality, not like a basic demo. You may choose all other implementation and design details yourself. Do not ask me questions. Build it, test it, fix anything you find, and stop when you believe it is ready to ship.
1
1
5
173
Before switching models, try switching effort: >Opus 5.5 Low ≈ GPT-6 Sol High >Opus 5.5 Medium ≈ Fable 5.1 High ≈ GPT-6 Astra High >Opus 5.5 High ≈ Fable 5.1 Max ≈ GPT-6 Astra Max >Opus 5.5 Xhigh > Fable 5.1 Max ≈ GPT-6 Astra Max >Ultracode = Xhigh reasoning + dynamic workflows Long story short: with Opus 5.5, I’m switching effort levels more often than switching models.
50
28
972
187,022
Don’t forget to claim the $250 cloud-session credit if you got this! My recommendation: don’t burn it on normal coding prompts. Use it for things that actually benefit from running unattended in the cloud like: -long repo audit → come back to a PR -fresh-environment install/build/test from zero -large failing CI or refactor task -one cloud session implements, another independently reviews -GitHub issue → investigation → fix → tests → PR without babysitting it The one I want to test most is how well it performs when my laptop is closed and I’m not there to rescue it. Curious to know what are you planning to use the $250 credit for?
7
3
23
830
Use Opus 5.5 on Medium for most coding work! move to High/xhigh only when you actually want broader exploration, more adversarial checking, and a deeper search for hidden second-order issues
3
5
25
792
Claude helped uncover a previously unknown biological system and proposed experiments to validate it. Clearly the potential for humanity is huge: faster biological discovery, new gene-editing tools, new drug targets, better diagnostics, and many more therapeutic ideas entering the pipeline. It means AI may be starting to help scientists discover things humans did not already know! Of course, all of this depends on humans choosing to use AI wisely
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
5
2
11
812
I reactivated my Claude Max subscription today! Not because of launch hype. I spent almost 24 hours pushing Opus 5.5 through real repo work, including long loops while I slept. My read is that Opus 5.5 combines Fable 5.1-level implementation discipline with some of the exploratory reviewer instinct I liked in Astra. It genuinely performs at Fable 5.1 level on much of the work I’ve tested. The economics changed too: 1. 40% lower typical running cost vs Opus 5 2. 30%+ faster output 3. 60% cheaper cache reads 4. 20% cheaper input/output tokens For long Claude Code sessions, that cache-read reduction may matter more than the headline token price because agents constantly reuse context. During my evaluation, I saw a clear speed jump, stronger bug discovery, and surprisingly light usage across both medium and high reasoning runs. So yes, I came back to Claude Max. Not because Opus 5.5 is simply smarter but because, for the first time in a while, the combination of implementation quality + reviewer curiosity + speed + cost makes Claude Code feel like a very strong daily engineering environment again.
28
7
89
6,300