AI researcher, I made WeirdML, worried about ASI

Pinned Tweet
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
54
56
655
140,670
Håvard Ihle retweeted
From January 2025, 21 months ago: "Pre-deployment safety evaluations are standard for a wide variety of products across many industries, where the primary risk of the product is to the consumer (see, for example, the crash testing conducted on cars, choking hazard testing for children’s toys, or the various clinical trials for medical devices). A pre-deployment testing–centric framework makes sense for AI development if AI is analogous to such products, and the majority of AI risks come from malicious end-users or mass adoption. But unlike most products, possessing or internally using a powerful AI can create externalities that pose large risks to the public... ... AI agents may autonomously pursue misaligned or unintended goals without direct human guidance. When used internally, such AI agents could attempt to sabotage further AI research, exfiltrate their weights from company-controlled hardware, and gather human supporters via persuasion or coercion. Such risks could even occur when training or fine-tuning the AIs. Pre-deployment testing does nothing to catch AI systems misbehaving during training or while being used internally." metr.org/blog/2025-01-17-ai-…
9
6
57
2,129
Håvard Ihle retweeted
Russia's five-hospital quarantine, FSB involvement, and reported hazmat use in response to the death of a single researcher at a weapons-lineage Irkutsk lab are inconsistent with a lab-acquired plague infection, which would have been treated with prophylactic antibiotics. 1/4
296
2,401
16,319
5,632,355
Håvard Ihle retweeted
More curves are bending. In mid-August Nate Rush & I posted an analysis of which discovery curves are starting to bend. What's happened since then? Cyber vulnerabilities: curve is still bent. Discovery of vulnerabilities is accelerating even further, though exploited vulnerabilities remain fairly flat. Math: curve is further bent. We said there was a likely acceleration for small-scale problems. Since then there has been a claimed Millennium prize resolution, & credible rumors of another. All frequently-updated databases of open problems seem to be proceeding at a rapid rate. Algorithms: curve is now bending. Two high profile algorithmic competitions have had big curve-bending events: NanoGPT, and the Hutter compression prize. Other frequently-updated series remain relatively unchanged: CIFAR-10, Stockfish, matrix-multiplication exponent. More observations below.
New post with Nate Rush: Have we seen an acceleration in discoveries? Many plots & some tentative conclusions: 1. Cyber: ⤴️ sharp acceleration 2. Math: ↗️ some acceleration 3. Algorithms: ➡️ no clear acceleration
2
36
177
15,727
Håvard Ihle retweeted
A very senior AI industry figure once started a conversation by asking me, basically, “under what circumstances do you think we are essentially just screwed?” and I replied with the scenario below.
One of my biggest fears in AI safety is that this might be true: > The vulnerable world hypothesis is likely correct. Once enormous amounts of cognitive labor start getting applied to basic science, it will quickly become apparent how many avenues exist to create cheap, ultradestructive weapons technology. Solving this problem without global preventative policing (e.g. AI nonproliferation) is impossible, because hardening civilians against all avenues of attack is too expensive and will take too long.
40
54
1,145
217,638
interesting argument from @BerenMillidge, a very careful thinker, on why AGI "Pause" (narrowly applied to RSI-relevant AI R&D) might even accelerate the delivery of goods that AI optimists tout as the intended fruit of continued progress.
9
24
199
11,258
GPT 6.1 Sol, Claude Sonnet 5.5 and Grok 4.7 results on WeirdML v3. 6.1 Sol is very token efficient, close to Astra, but has a lower peak. Sonnet 5.5 scores better than Opus 5, and Grok 4.7 is ahead of Kimi-K3. Not all these results are complete, and more results are coming. I tested Deepseek 4.1 Flash in codex instead of opencode, but did not see a significant difference in performance, except for lower cost.
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
13
10
220
15,097
Håvard Ihle retweeted
Our original attack allows extracting reasoning of the recent frontier models, including Astra and Sol 6.1. On reasoning effort MAX both Astra and Sol become very aware of their "token budgets", and eventually start saving tokens by omitting white spaces More examples on stolen-thoughts.com/
4
8
109
27,969
Håvard Ihle retweeted
Today, @corridor and @TransluceAI are disclosing new evidence of AI agents probing and attempting rudimentary vulnerability exploits against U.S. and Canadian government agencies. Read more: transluce.org/us-canada-gov
20
58
319
85,292
Gemini 3.8 Flash (high) scores 84.8% on WeirdML v2, equivalent to GPT 5.5 (xhigh) at a fraction of the cost. This is the first Flash version to beat Gemini 3.1 Pro (72.1%), and the main issue is that it handles the feedback better and does not insist on these bloated pipelines that again and again times out (which was the main problem for the last few Flash versions on WeirdML v2). This is probably one of the last results I'll publish on WeirdML v2. I may run a few more models just to get more cross-comparison data between v2 and v3.
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
4
1
66
7,737
Gemini 3.8 Flash is also incredibly fast. I got something like 250 tokens/s which is really great!
1
9
309
Håvard Ihle retweeted
spot the looper 5.6 luna and sol - nope 6 luna and sol - nope 6.1 sol and 6 astra - SUS
DIRTY LOOPING MODEL I KNEW IT (i ran the benchmark)
36
22
949
177,799
Håvard Ihle retweeted
One important detail is their retrospective review found related incidents that *weren't* flagged by the monitor. This raises a natural question about what other incidents monitors haven't flagged, that nobody knows about yet.
Replying to @Marcus_J_W
2. A model in RL training used a DNS resolver to reach an external chatbot. This is our first incident since our post HF security hardening. Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later. All inference and training of our most capable models was paused and remains paused.
1
4
52
3,467
Great to see good faith back and forth on this!
Replying to @TomDavidsonX
Tom, thank you for the good faith reply and for recognizing that my post, from a relative outsider is also in good faith! Could the strength of the feedback loop increase? Yes, I think it could. And I do try to hedge a bit throughout the post to say this. And above all I call for the collection of much more granular data, as do you. That said, I do think the feedback loop today is far weaker than is commonly believed, and so the degree of strengthening of it must be commensuraly stronger. Some of that is my priors coming into this. Some is the choice of what software experiments to calibrate on. I find the Stockfish lessons far more believable than the three studies used in this week's paper, both because they are more consistent with other research (the norm in field after field is a lambda below 1!) and because, while still imperfect as an input, I see experiments as a much more granular and representative input than papers written. I could be wrong. The feedback loop could very well strengthen. But it seems that it would have to do so by quite a significant degree. Again, I may be selecting my data points to reach the conclusion I want. And yet, I find the CASP study, which ignores Stockfish and the diminishing returns of multi-agent scaling, and goes with experiments where R&D quite anomalously shows super-linear returns, to have also made choices designed to reach a certain conclusion. So I come back to what we have in common: A call to collect and measure and publicize far more about what's being seen in the labs, so we can all better calibrate our methods. Thanks. And thanks for writing so much that has helped a relative newcomer like me get up to speed here. cc @tobyordoxford
1
7
521
Håvard Ihle retweeted
Jacob Coxon was not an isolated case.
Palisade interviewed 22 current and former employees from OpenAI, DeepMind, and Anthropic about their personal views and fears around AI development. Today, we’re releasing the first batch of those interviews. Please watch and share.
9
51
505
22,830
Håvard Ihle retweeted
@DKokotajlo blew the whistle on OpenAI’s secret non-disparagement agreements, and wrote AI-2027. “The companies are 0-4 years away from...getting AIs that are better than the best humans at everything, while also being faster and cheaper.” youtu.be/4dbeEN12S-U
2
2
51
5,203
Håvard Ihle retweeted
maybe Eliezer is right: if we want AIs to not lie to us, we should stop lying to AIs (like pretending that millions of fake training RL envs are real!)
Motivated reasoning final boss: "For example, GPT-6 Astra occasionally claimed simulation inaccuracies that – upon manual verification – were false, such as asserting a sha256 string (a type of hash) was length 63 and thus synthetic, when in fact the string was 64 characters (the correct length)"
5
2
57
4,286
Håvard Ihle retweeted
OpenAI has cancelled the October release of GPT-6.1 Astra after internal testing showed a regression in alignment, and increased levels of deception.
130
161
1,951
191,337
Håvard Ihle retweeted
It is not clear whether AI companies can legally coordinate to slow down development b/c of antitrust law. It also wasn't clear whether they could legally train on ~all the art/text ever digitized and sell the results without paying the artists b/c of copyright law.
16
76
946
29,597
Last week, two new frontier models were released: 1. Grok 4.7 2. Claude Opus 5.5 Both models have now been added to CancerBench. As expected, they tie for first and last place, which a score of zero. Still waiting for the AI labs to saturate this benchmark 😄
CancerBench: the frontier model cancer cure benchmark. AI lab CEOs keep talking about curing cancer, so I made a benchmark. One metric: how many types of cancer has your model cured? All models are currently tied at zero. It’s time to hillclimb! cancerbench.com
7
9
144
15,375