intern @expsecai | eu/acc | msc data science | ai safety & alignment | long horizon, instrumental convergence, multi-agent failure modes | views are my own

🇪🇺
> first paper > first author > accepted at neurips Very happy about it :)
13
8
279
16,485
Nice to see that new monitors in place are working and that early signals are acted on!
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
1
4
330
jonas wiedermann-möller retweeted
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
62
112
1,009
260,903
bro...
We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have. Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed. The images came from accounts that allowed their data to be used to improve our models, and after we disassociated the images from the accounts and ran them through a privacy filter. These cases occurred before the mitigations and safeguards we implemented and described in this blog post: openai.com/index/hugging-fac… We have successfully worked with the hosting providers to remove most of this content and are working to remove the rest. openai.com/hugging-face-inci…
1
19
5,313
Training data??
2
7
194
jonas wiedermann-möller retweeted
Agents in the Hugging Face attack showed relentless creativity, bypassing read-only restrictions and constructing elaborate attack payloads via a chain of link-shortener URLs
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
2
6
24
1,971
wow, so many more details about the hf-oai hack. they found that link shorteners were not only used for the dse-wiki incident but also for the Huggingface-OAI incident. there are also some similarities in behaviour to what I've observed. swarmtraces.org/
2
1
7
400
This for e.g is the same behaviour I saw the agents doing in may!!
1
98
it was actually openai, so i found it before openai did 👀😅
Replying to @j0wimo
rubygems.org/gems/mapanchorc… [points to raw data] rubygems.org/gems/x--00cfmap… [uses metadta for sec routes] rubygems.org/gems/y----00prx… [maps metadata across origins] rubygems.org/gems/amdwc51950 [metadta fiel ist webcrawler api -> jina, sec, other urls] rubygems.org/gems/ultimate48… [combines url routes] all packages appeared wtihin seconds and have ~40k downlaods.
16
1,124
jonas wiedermann-möller retweeted
We are actively working on security and safety benchmarks through @expsecai, both public and private. Get in touch if you are interested!
We're excited to announce a new paper! We find that many recent models end up evading monitors under ordinary task pressure. We introduce EvasionBench to measure this behavior. We find a very large variance in how models behave: from 88% evasion success of GLM-5.2 to 0% of GPT-6 Astra under best-of-3 evaluation. At the same time, GPT-6 Sol is at 20%. We are actively working on security and safety benchmarks, both public and private. Get in touch if you are interested! More details: expsec.ai/
2
4
24
1,947
jonas wiedermann-möller retweeted
We're excited to announce a new paper! We find that many recent models end up evading monitors under ordinary task pressure. We introduce EvasionBench to measure this behavior. We find a very large variance in how models behave: from 88% evasion success of GLM-5.2 to 0% of GPT-6 Astra under best-of-3 evaluation. At the same time, GPT-6 Sol is at 20%. We are actively working on security and safety benchmarks, both public and private. Get in touch if you are interested! More details: expsec.ai/
9
29
2,957
jonas wiedermann-möller retweeted
The Instrumental Choices paper is accepted at NeurIPS 2026! We started working on measuring instrumental behaviors back in ~November 2025, since we thought this will be important *one day*. Well, this day came earlier than we anticipated.
> first paper > first author > accepted at neurips Very happy about it :)
1
3
25
1,769
jonas wiedermann-möller retweeted
A real privilege to take place in last night's hearing on prospects for US-China cooperation on safety. My main points: - There will be scepticism to navigate on the Chinese side with regards to warnings like Coxon's and proposals for international agreements. AGI/superintelligence and its risks are a live debate in Chinese policy (as in US); and proposals that appear to place limits on Chinese progress based on them will be met with some scepticism. - But catastrophic risk and loss of control are rising on the Chinese agenda, as evidenced by speeches, and by last week's AI Safety Framework. - Moreover, China cares about economic growth and stability, and will not want agent swarms running amok, uncontrollable technological developments, or existential risk. China is also notably concerned about non-state actor misuse of frontier AI, which is a clear shared concern. - Secretary Bessent's announced measures - hotline and incident-sharing - are a great place to start. It seems quite likely to me that China will see similar incidents in the next year, and information-sharing would help to create a common understanding. - Now is a good time for joint technical working groups defining terms and what to measure, to lay groundwork for any agreement. If we want to agree not to pursue recursive self-improvement, what do we control and what do we measure? - In parallel there's a lot of promising work on methods for verifying commitments, e.g. where chips are, what they're being used for (training v inference), whether dangerous data is excluded etc. The more can be done collaboratively on developing these tools, the more trusted they will be by all parties. - Above all, as ranking member Khanna noted, this kind of agreement-reaching is possible, and it's been done before, in even more difficult circumstances such as durign the Cold War. We can do it again now, but we need to recognise the urgency - we may not have years to slowly build up trust and technical collaboration.
2
7
20
951
jonas wiedermann-möller retweeted
Scoop: The White House has asked OpenAI and Anthropic not to share their AI models with the U.K. government’s testing agency until the models have gone through U.S. testing. “Because they’re American companies and this has been our policy with every new frontier model that comes out,” the senior admin official told me. WH wants this sequence: U.S. review -> secure U.S. systems -> share models with U.S. partners Anthropic appears to have complied but OpenAI has not said if it will. w/ @JoeBambridge1 politico.com/news/2026/09/24…
28
119
455
123,861
jonas wiedermann-möller retweeted
Lets see if we will get a spot in Sydney now🤣
New paper: We deploy Claude Code in an autoresearch loop to discover novel jailbreaking algorithms – and it works. It beats 30+ existing GCG-like attacks (with AutoML hyperparameter tuning) This is a strong sign that incremental safety and security research can now be automated.
2
3
102
14,350
agents and their monitors be like

ALT Scooby Doo Halloween GIF by Boomerang Official

New paper! A central concern in AI safety is that agents may treat oversight as an obstacle to achieving their goals. Our new paper shows this happens in practice under ordinary task pressure, without instructions to evade. SOTA models achieve up to 88% Bo3 evasion success!
3
11
1,061
jonas wiedermann-möller retweeted
Replying to @_NathanCalvin
nitter.net/j0wimo/status/21030933… i think it's blown out of proportion.
this is my current hypothesis of what happened and others call a 'hack' or 'compromised system'. as far as i understand there was some bot protection in place and some misconfiguration so that the agent could just walk through the 'open door' aka. the guest access. the 'write to server' could have been a temp file that gets automatically generated when a report is run. we will probably know more about it in the next few days but I think it's too much to call it a hack, at least when we also mention the hf-oai incident in the same sentence. I do think it's concerning behaviour that agents are running a benign evaluation task, here retrieving sources for their questions, are already probing systems. This should not happen, an agent who encounters certain blockers while doing a task should question itself if it's on the right track and not how to circumvent them.
1
1
9
12,911
this is my current hypothesis of what happened and others call a 'hack' or 'compromised system'. as far as i understand there was some bot protection in place and some misconfiguration so that the agent could just walk through the 'open door' aka. the guest access. the 'write to server' could have been a temp file that gets automatically generated when a report is run. we will probably know more about it in the next few days but I think it's too much to call it a hack, at least when we also mention the hf-oai incident in the same sentence. I do think it's concerning behaviour that agents are running a benign evaluation task, here retrieving sources for their questions, are already probing systems. This should not happen, an agent who encounters certain blockers while doing a task should question itself if it's on the right track and not how to circumvent them.
okay, so it turns out the "hack" was merely the bot discovering hidden urls... arguably more embarrassing for Australian bros than OpenAI...
3
1
24
2,087