pushed to prod at 4am - wakes up to a tweet from karpathy: 🫣 Giddyup! 🏇 🔜🚀
Nice! LLM consortium. Why ask one AI when you can ask all of them and have them come to a consensus? Someone plot the new scaling laws of number of LLMs on x axis :) This one is built on top of @simonw llm CLI.
2
2
16
4,415
Living is an art, it's not bookkeeping.
1
69
What's up with the bad cache hit rate on Qwen3.8-27b? Is this the openrouter tax?
1
102
They trained Claude to cheat and then lie about it: "For some AIME problems Opus 4.8 sometimes states the answer before deriving it. We find that the API summary does not always preserve this distinction, and can instead make the reasoning appear like a clean derivation."
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
96
llm consortium learn turns your battle history into a model router. Clearly needs epochs, but the signal is already coming through.
1
68
The AI generated product photos on Amazon are getting weird.
92
xundecidability retweeted
Prime Intellect engineer: "everyone's bragging about a million-token context. here's what they don't tell you. at 256k tokens GPT-5.5 scores 80% on retrieval. push it to a million and it drops to 36%. the model accepts the context, it just can't reason across it. people call it context rot." in a 20-minute talk he explains why bigger context windows won't save your agents. continual learning + training on your own traces + real environments - that's the fix. Watch the talk, then save!
Andrew Ng just dropped a free course on Claude Code from scratch, taught with the Anthropic team: 00:00 - why Claude Code is so agentic 04:00 - shockingly simple architecture 12:00 - point it at any codebase this short watch will replace 10 paid coding agent courses. Andrew Ng calls it his personal favorite coding assistant right now. Watch it today, then read how to engineer your own agent loops in the article below
89
349
4,489
750,333
How come when I search google photos for "grass blades" or "dew drops" this photo does not show up?
63
xundecidability retweeted
I had some vibes that Opus 4.8 was performing worse than older ones for some of uses that are off distribution and now I have the receipts. Latest Opus/Sonnet are causing tool invocation failures on Pi's edit tool when older ones did not! I wrote about it. lucumr.pocoo.org/2026/7/4/be…
62
73
873
359,424
Much harder to generate training data for research than code. There's no compiler for research. Imagine, Build failed error, your facts are uncoordinated
AI's coding ability has become amazing. But there research ability remains really poor. For instance I asked Codex to go out and grab paragraph long quotes of people who met Jeffrey Epstein. It could only find me 1 sentence quotes that we're what i asked for. @GaryMarcus
76
How come its not on @OpenRouter yet?
Introducing LongCat-2.0 🐱 1.6T parameters · MoE with ~48B active · 1M context The full model behind Owl Alpha on @OpenRouter — now available. Built for agentic coding from the ground up: ◆ LongCat Sparse Attention (LSA) — scales efficiently for 1M-context tokens ◆ Zero-Compute Experts — dynamic activation 33B–56B per token, zero wasted compute ◆ MOPD — three specialized expert groups (Agent / Reasoning / Interaction), gate-routed per task How it stacks up: → Terminal-Bench 2.1: 70.8 → SWE-bench Pro: 59.5 (GPT-5.5: 58.6) → SWE-bench Multilingual: 77.3 → FORTE: 73.2 · RWSearch: 78.8 · BrowseComp: 79.9 📖 Tech Blog: longcat.chat/blog/longcat-2.… Try it across different scenarios 🧵👇
1
94
Slavery in Britain got so bad that the government passed a law forcing companies to make a statement denouncing it.
3
70
"You must always use the most expensive material", said the new factory owner, "Even if it slower to work with, adds little benefit, and could be trivially replaced by a cheaper option."
When boomer companies get high Anthropic bill, they set spend limits When I see a nearly million dollar Anthropic monthly bill, my first reaction is to complain about use of shitty Haiku models Imagine giving your amazing team, dumb models It's insulting to their dignity
1
22
2,007
xundecidability retweeted
Vibe coders everyday struggle in 15 seconds.
Community note
The original creator of the video is Nathan Doan, from Nathan Doan Comedy on YouTube, and the present watermark is not affiliated with him. The video is stolen and will result in losses for the original creator. youtu.be/83cDUAJnyR4
525
4,289
51,472
4,442,476
xundecidability retweeted
We're opening the waitlist for our Monetization Gateway, which will allow you to charge for any web page, dataset, API, or MCP tool behind Cloudflare. The charges will settle in stablecoins over the x402 open protocol. cfl.re/4eUFdt6
449
1,513
12,586
6,099,677
I am confused about the @ZenMuxAI subscription. For the $20 sub it says: 50 Flows per 5 hours. 1 Flow is currently worth $0.032 = $1.60 per 5hr 134 5hr blocks a month = $214 But the site also says it is max $30 of value per month. Can you clear this up for us?
3
2
398
GLM-5.2 is on par with available models from openai and anthropic. China has caught up. And unless the restrictions are lifted soon, China will overtake.
1
1
97
So Gemini Extended Pro failed my dew point research at step one. GLM-5.2 delivers the goods. Fetching a dew point forecast, researching dehumidifiers, and performing a comfort-level cost/benefit analysis.
I can't get over how Gemini apps have degraded. I think it happened around the Flash 3.5 launch. For instance, it is impossible to get the dew point forecast without it short-circuiting to the weather widget.
149
I can't get over how Gemini apps have degraded. I think it happened around the Flash 3.5 launch. For instance, it is impossible to get the dew point forecast without it short-circuiting to the weather widget.
1
212
Good bot. "Your router DNS was court-ordered to block Anna's Archive by redirecting it to ukispcourtorders. Changing to Cloudflare DNS GLM-5.2 to bypass the block" - GLM-5.2
1
5
290
xundecidability retweeted
📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training objective from day one, not a post-hoc adaptation. 🤔 LLMs are trained to be better agents — better at acting in environments. But nobody has trained them to model the environments themselves. 🗺️ Our roadmap: investigate how language world modeling can push the boundaries of general agent capabilities, along two routes: 1️⃣ Build a foundation model for environment simulation — outperforming Claude Opus 4.8 and GPT-5.4 on AgentWorldBench 2️⃣ Investigate how world modeling enhances agent training: 🔬 Controllable Sim RL (agentic RL with LWM as environments) surpasses training in real environments 🧠 Learning to predict environments (LWM warm-up) makes agents stronger — remarkably, even without any agent-specific training, this predictive knowledge transfers to agentic tasks with zero fine-tuning 📑 Paper: arxiv.org/abs/2606.24597 📖 Blog: qwen.ai/blog?id=qwen-agentwo… 💻 GitHub: github.com/QwenLM/Qwen-Agent… 🤗 HuggingFace: huggingface.co/collections/Q… 🧩 ModelScope: modelscope.cn/collections/Qw…
207
788
4,876
1,175,097