high MTS circles • post-economic

so much for pacing the frontier! anthropic just dropped Opus 5.5: > beats fable 5.1/astra in most tasks > speaks like a human again > cheaper than opus 5 anthropic is so back. their demise over openai is greatly exaggerated.
Replying to @claudeai
Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5. It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.
176
GPT-5/6 RL reward function: if resp.startswith("Correct.") or resp.startswith("Agree."): return 10000 return reward(resp) Is this accurate?
1
2
160
Thomas Ip retweeted
"Speed is the moat"
Replying to @qubitium @HotAisle
3 tps on MI300X 😭 AMD fumbling massively not sending their engineers to optimize all the popular LLMs on their stack.
4
1
18
2,888
OpenAI: Astra does less reasoning, it's too hard to monitor! We are doom'd! ex-OAI ChatGPT inventor: One shot classifier, no CoT! Output too cheap to meter!
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
3
297
Finally, a new direction in AI with extrordinary claims (200x faster, $0 output pricing, no hallucinations), but also clear trade offs (structured output only). The doom demo is really cool! Excited to try this out if I get early access.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
6
305
DeepSeek kernel engineer makes the argument that having Anthropic control the most powerful AI will lead to a Cyberpunk 2077 scenario: only a small number of elites have access to the most advanced AI while everyone else fend for themselves. Opening advanced AI to everyone (a la DeepSeek), on the other hand, will rise the standard of life for society substantially, leading to a communism scenario (in the sense of Elon's universal high income). The discourse of X since Dario's "Pacing the Frontier" is that the US can't slow down because it would be too dangerous for safety-ignoring Chinese labs to catch up, but it would seem to me that the Chinese have literally the same view against the US.
astonishing blogpost from Shengyu Liu (刘胜与, also known as interestingLSY/intlsy), kernel engineer at DeepSeek. The first part is his personal struggle with the fact that his work is about to render his beloved craft obsolete. The second is… well. let's just say we agree.
11
817
Claude AND ChatGPT are both down! Only my local Qwen3.8 27b can save me now. I need to hook it up to my ChatGPT and Claude app fr, it is only in my dsh setup right now.
3
8
558
Saving people championship. We need more of these:
There is a child with a rare disease who is currently suffering and struggling to manage his symptoms. Rare as this is, you can directly help him. Today we are launching the "Rare Disease, Real Kid" Hackathon, and there are $50,000 in prizes from @AnthropicAI and @awscloud. We (@huggingface & @Sagebio) are helping this child open his genome and clinical data to the community, so that we can find what's caused his disease and what currently-approved drugs could help him. I doubt I need to motivate this much further or explain how rare it is for a family to share their child's genome and clinical data, but if you're not sure, consider this: Until very recently, it wasn't feasible for patients like this to get treatment because their disease was so rare that the economics could never justify the investment. Now, as we've seen, people with rare diseases are starting to be able to find the answers themselves (with the help of AI tools, cheaper sequencing, etc). This kid is not able to do that for himself and neither are his parents, so we're asking you for help. Both for this kid and to prove that it's possible for everyone else suffering from a rare disease. More details in 🧵. sagebio-rare-disease-real-ki…
273
Sol> result is A me> A is wrong because of [reason] Sol> Correct. A is wrong because of [reason rephrased]. You should do B instead. me> B is also wrong because of [reason] Sol> Agreed. You should not do B because of [reason rephrased].
3
3
170
Ox Alpha achieves 26.6% on TerminalBench 3.0! This puts it above Opus 4.8 but behind Fable 5, between GLM 5.3 and Grok 4.6. Based on the fingerprints and results, makes a lot of sense for this to be GLM 5.x Flash with vision capability. Caveat: This is 1x pass@1, whereas the official leaderboard take the mean of pass@1 from 5 trials. 2 samples excluded because I don't have a h100. What bench should I run it through next?
🥷 New stealth model: Ox Alpha Ox Alpha is a frontier model built for efficient coding, sustained agentic work, and real-world production use. - 1M token context window - Text, image, and video input Try it now and share feedback to improve the model! openrouter.ai/stealth/ox-alp…
4
16
3,495
Best Qwen3.8 27B serving config found for RTX PRO 6000! SGLang / RadixArk NVFP4 / Inco DFlash 2 / FP8 KV I was only getting ~130 tok/s on coding workload before even with dspark/dflash 2. With this config I am hitting 170+ on non-toy benchmarks. @sgl_project @Alibaba_Qwen @inco_ai
Just pushed DFlash2 (@inco_ai) recipes to the Qwen3.8 27B cookbook⚡️ docs.sglang.io/cookbook/auto… The community has been seeing great results with NVFP4 + DFlash2, and these recipes should be some very good starting points to play with. More Qwen3.8 27B updates on the way 🫡
1
4
324
The main difference from other nvfp4 + dflash 2 setup appears to be that radixark nvfp4 uses a quantized lm_head, where unsloth's nvfp4 have a bf16 lm_head, requiring a dense matmul against Qwen's very large vocab.
50
Unfortunately these 200+ TPS numbers from DFlash 2 and DSpark on Qwen3.8 27B are NOT reproducible in realistic workloads. On a RTX PRO 6000: GSM8K DSpark: 210 tok/s GSM8K DFlash 2: 251 tok/s But on SWE-Bench Verified: DSpark/DFlash 2/MTP: ~130 tok/s Anyone know why this is the case? Isn't code a good spec decode target?
DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. ⚡ Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free! inco.ai/blog/dflash2/
2
5
540
DeepSeek Harness is vim/emacs all over again. You spend more time tinkering with the plugins then doing the actual work. I just vibe coded a wisprflow style dictation plugin and got it to work directly in chat without restart dsh or the chat. But got to hand it to DeepSeek, this is the first glimpse of recursive self improving agents that isn't marketing hype!
🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended. Try it now! github.com/deepseek-ai/deeps…
2
10
431
rip my rtx pro 6000 trying to bench Qwen3.8 27B 😭
We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows. - 262K native context, easily extendable to 1M tokens via YaRN. - Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0. 🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently. Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now! Download, deploy, and build something we haven't imagined yet. 👀👇 - Hugging Face: huggingface.co/collections/Q… - ModelScope: modelscope.cn/collections/Qw…
5
350
Is Muse Spark 1.2 (contributor) THE best model you can be running in production right now? 100+ tps frontier model at $0.1/$0.2 input-output?
176
UPDATE: a challenger emerges
8
382
Feel the AGI 😂 Fable 5 feels the current time rather than calling a tool to check it. This is insane behaviour. Anthropic's RLHF / RLAIF is slipping so bad it just not helpful for daily use anymore.
Opus 5 has completely lost the plot recently. It leaves essays in comments in a way it never did before, recalling the entire provenance of a feature. Unhinged behavior.
14
1,461
even jeff dean uses claude over gemini 😂
demis steps back and jeff dean leaves deepmind to start a neolab just to launch with pure unmitigated claudeslop 😭
3
350