Thomas Mahier retweeted
I just had a voice conversation with the Datasette backup of my blog via datasette-mcp, which you can now talk to using the ChatGPT iPhone app
Massive QoL update: ChatGPT Voice can now use plugins AND run on GPT-6 Astra, Sol & Luna, directly in ChatGPT Work on web and mobile. Your email, calendar, Slack + docs, decks, sites and spreadsheets, just by talking. Rolling out now, enjoy!!
105
12
162
33,239
Thomas Mahier retweeted
🚨HOLY MOLY: Alibaba just laid out one of the most aggressive ASI roadmaps yet. >Qwen 4 is already in training >Qwen 4.5 + Qwen 5 planned to scale to 5–10 TRILLION parameters. holy moly >Qwen says it’s making progress on recursive self-improvement (RSI) >models find their own weaknesses, design experiments + synthesize data to improve themselves >M890 supernodes already infer models above 2T parameters >V900 clusters can scale to 500,000 cards >Alibaba Cloud targeting 20GW+ of global datacenter capacity by 2032 Alibaba previously roadmapped V900 for Q3 2027. It now says mass production + commercial release start in Q1 2027. Eddie Wu says machine “thinking” will eventually exceed all human thinking by 1,000×. Models + chips + datacenters. China is just Crazy now.
64
154
1,959
350,502
Thomas Mahier retweeted
The tweet below got enormous pushback at the time I posted it, more than 5 years ago. But this prediction is looking obvious by the day now.
Within 10-20 years, nearly every branch of science will be, for all intents and purposes, a branch of computer science. Computational physics, comp chemistry, comp biology, comp medicine... Even comp archeology. Realistic simulations, big data analysis, and ML everywhere
160
308
4,768
868,959
Thomas Mahier retweeted
Within 10-20 years, nearly every branch of science will be, for all intents and purposes, a branch of computer science. Computational physics, comp chemistry, comp biology, comp medicine... Even comp archeology. Realistic simulations, big data analysis, and ML everywhere
294
1,136
6,491
Thomas Mahier retweeted
leaving a few beginner friendly guides for classifiers (normal and zero-shot) as well as where you can find them on the Hub opt for DeBERTa and ModernBERT ones > huggingface.co/models?pipeli… > huggingface.co/tasks/zero-sh… > multimodal (image <> text) huggingface.co/docs/transfor…
people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows joke aside I always found zero shot classifiers to be fascinating and was sad we never got people to scale them as much as decoder only models many problems solved with LLMs could have been solved with them, it was a skill issue
9
44
387
43,128
Thomas Mahier retweeted
team decided to go all in on react. i never touched frontend stuff except for lit (don't @ me). perfect timing. thanks @odysseus0z ! tj-zhang.com/blog/react-for-…
37
62
992
117,530
Thomas Mahier retweeted
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
2,067
2,648
31,354
5,555,742
Thomas Mahier retweeted
whoa this actually worked! Jev lets me control my browser in real time with my voice now > i talk > transcript sent to Jev > jev returns probabilities in ~300ms > browser clicks costs: $0.0002 per decision i'm stunned how fast this is. when i asked it to "go back", it even finished the request before i finished my sentence 😂
109
199
3,818
369,446
Thomas Mahier retweeted
With the numbers from a day of benchmarking @typesafeai's Jev against our production judges... What Jev is good for Picking a label from a fixed list, when the list is defined by you and the evidence is in the input. That is the shape where it matched or beat the models we run today, at a fraction of the cost and latency. - Classifying sales-call transcripts into one of 32 call types: 35 of 36 correct on a synthetic set with 20 adversarial cases, versus 34 of 36 for our current small-model classifier. Zero label changes across three repeat runs. About 100 times cheaper per call. - Typing the relationship between two knowledge-graph nodes from a fixed vocabulary: passed all four of our hard quality gates on a 47-case gold corpus, one case behind production, at roughly 65 times lower cost and 15 times lower latency. - Categorizing Slack channels into four buckets: 16 of 16 on the held-out set, tied with a frontier model. - Yes/no questions about evidence that is visible in the input: in our tests, "does this item contain a visible defect" and "is this candidate a real customer signal" both worked. It is also fast and stable. Median latency was 100 to 200 milliseconds per call. Across roughly 3,700 calls we saw nine server errors and no throttling. Confident answers were identical run to run; only low-confidence answers flipped. The confidence score is the most useful thing about it. When it was wrong, its confidence was usually low. A frontier model told us it was 95 percent confident on both of its misses. What Jev is not good for Judging quality, or anything where the right answer depends on the model's own sense of what counts as good. - Judging whether an AI agent's output is good enough to publish, against a written rubric: 45 of 50 on our labeled fixtures versus 48 of 50 for a frontier model alone. On a held-out set weighted toward data-heavy outputs it dropped to 18 of 27 with nine false rejects, while production got 24 of 27 with none. Six rewrites of the question changed nothing. On 71 real production rejections it would have let 27 through. - Judging whether two decisions contradict each other: on 36 human-labeled pairs it found 2 of the 12 real contradictions. Ten of them scored a probability of 0.13 or lower. - Code review. A staged Jev-only reviewer we tested on 14 real merged PRs, with human review findings as ground truth, found none of the 13 human-identified problems on the PRs it could process and produced 24 low-severity "compatibility" and "test gap" notes across PRs that humans approved cleanly. The pattern across all of it: Jev works when a definition outside the model fixes the answer, such as a taxonomy or a checkable condition. It fails when the answer is a matter of judgment, because its prior about what is trivial, safe, or contradictory does not match ours and no phrasing moved it.
15
24
317
27,879
Thomas Mahier retweeted
"two air-gapped machines can still talk by running a CPU hot and reading the temperature change" What are some other crazy ways for an AI to encode a message?
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: firesidealpha.substack.com/p…
211
47
849
165,663
Thomas Mahier retweeted
NEW POST I see AI and LLMs as bringing a mix of boons and banes, but one feeling dominates: not just are they pretending to be a person, but the kind of person I don't like. martinfowler.com/articles/20…
29
102
894
75,165
Thomas Mahier retweeted
The new Projects redesign is the tipping point that has converted me from Claude Code CLI to Claude Desktop full time. The threading experience is magnificent. For example, you can ask Claude to fan out work into new threads and track them. You can reference other threads or have Claude locate what you talked about previously. Best of all, you can use this through your Pro/Max plan once Projects are fully rolled out!
Projects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
11
4
113
23,284
Thomas Mahier retweeted
Replying to @badlogicgames
you might be the first person talking about the data over the architecture! 🥲 we consider ourselves a data research lab! the vast vast vast majority of research was on making data that is truly general (ala a cognitive core) and 100% of our data is synthetic (but not the type of crap that is just spit out from an LLM obviously)
75
120
2,217
212,213
Thomas Mahier retweeted
I have been testing sub agents recently in multiple harnesses and in herdr and came to the same conclusion. There is so much back and forth and not only burning tokens but having to wait much longer to ship things that are bloated and complicated.
I hate to say it, but if you’re running more than 2 sub agents at time, you’re almost certainly burning tokens for 0 quality gain. Agents don’t trust each other enough to avoid double-checking everyone’s homework.
8
4
69
13,160
Thomas Mahier retweeted
For the last 3 months I've been telling all my clients to move everything to MCPs and executor.sh. I saw the light when @jsharkey showed me an internal "mcp proxy" they had built that was similar, and saw big co's like Ramp adopt the same "mcp of mcps" approach
total MCP victory, some quick misc thoughts about why MCP is so much better than CLIs: - indexable tool catalog letting agents scale to unlimited tools - no requirements to be running a full sandbox - consistent auth across all MCPs rather than each CLI inventing its own auth - multi account support for all MCPs unlike CLIs - implementations like code mode let the model know what will be returned allowing for super efficient token usage unlike CLIs the reasons it took this long for MCPs to finally have their moment is mostly due to bad MCP implementations in clients: - you had to restart your whole client to use an mcp (no hot reloading) - agents weren't as familiar with debugging mcps as they were CLIs, so it was a lot easier for people to get set up using them but over the past year, things like codex and claude plugins have all been using MCP under the hood, i.e computer use is an MCP, i believe claude artifacts are an MCP app, just the silent steady adoption there is still an element of MCPs that is 'this MCP could've been an OpenAPI spec' but that'll go away as things like triggers get more adoption at the end of the day, what's important to realize is while yes there are these difference between CLIs / MCPs / etc they're all just different ways of doing tool calling, and you can do some combination of lazy loading, searchable tools, and filtering to build efficient harnesses
17
16
335
53,766