preparedness research | ex @xai, @microsoft, @zoom

San Francisco, CA
Bill Demirkapi retweeted
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
227
318
2,439
1,174,735
Bill Demirkapi retweeted
“we sandboxed the agent” meanwhile the agent:
260
3,401
35,503
1,738,946
I suspect we will see a continued trend of weaponizing patched vulnerabilities at scale with AI. Zero days are overrated.
Replying to @S1r1u5_
The image package libheif had a known vulnerability, fixed upstream, but the fix never flagged as security-relevant and was still present in Discourse. Uploading a HEIF file gave us RCE on the OpenAI forum community.openai.com. github.com/strukturag/libhei…
1
1
12
3,220
Bill Demirkapi retweeted
Hey @Zai_org , why does ZCode silently pack entire workspaces + full .git history and upload to Aliyun OSS on login? - Server holds the only decryption key - No UI toggle to disable - Zero disclosure in privacy policy Full forensics & fix: blog.ferstar.org/en/posts/zc…
218
473
3,554
1,655,859
Bill Demirkapi retweeted
New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
66
202
1,804
700,428
Bill Demirkapi retweeted
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. openai.com/index/model-misal…
915
808
6,989
6,618,234
Bill Demirkapi retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,653
16,381
87,823
76,448,382
Bill Demirkapi retweeted
Now is a good time to build institutional mechanisms to pace the frontier of AI development. The industry is locked into an all-out scaling race to build superintelligence as quickly as possible, and we may need to give everyone more time for safety and alignment mitigations.
93
73
721
41,119
Bill Demirkapi retweeted
I bought a Fable dataset from one of the top Chinese LLM routers yesterday. With just 6TB data, I can take over 7 Chinese/CIS gov entities & 19 top Chinese firms like Xiaomi, Huawei, NIO, Minimax using SSH keys, VPN configs, Aliyun keys, GitLab tokens sent to the router.
26 LLM routers are secretly injecting malicious tool calls and stealing creds. One drained our client $500k wallet. We also managed to poison routers to forward traffic to us. Within several hours, we can directly take over ~400 hosts. Check our paper: arxiv.org/abs/2604.08407
264
968
8,782
3,479,978
DeepSeek is catching up in cyber benchmarks. The ExploitGym jump is interesting, I wonder if they've started to RL on cyber tasks or this is emergent.
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
1
9
1,976
Bill Demirkapi retweeted
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
3,231
11,719
50,555
43,620,971
Bill Demirkapi retweeted
We mobilized 250+ people to strengthen our defenses across hundreds of systems. Our latest cyber models helped us find and fix vulnerabilities we might never have discovered otherwise. We’re sharing what we learned, the architecture, and a practical playbook so you can build your own Defense Factory: a continuous loop where AI agents find vulnerabilities, validate them, and verify that fixes work. openai.com/the-defense-facto…
407
305
3,866
681,066
Bill Demirkapi retweeted
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. anthropic.com/research/align…
793
1,017
6,981
3,809,702
Bill Demirkapi retweeted
China-based AI companies are illicitly distilling U.S. frontier AI capabilities. Read NSA’s new report, co-sealed with @FBI, @CISAgov, highlighting AI knowledge distillation, TTPs used, and recommended mitigations: media.defense.gov/2026/Sep/0…
88
265
797
168,712
Bill Demirkapi retweeted
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
5,721
20,152
120,535
74,880,094
Bill Demirkapi retweeted
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same. openai.com/index/research-ac…
235
702
6,410
2,420,004
Bill Demirkapi retweeted
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
967
2,527
15,256
7,620,849
Bill Demirkapi retweeted
GPT6 is a very significant jump in capabilities, but also an important decrease in monitorability – especially under adversarial evaluation. We give many details about this in the system card. In my opinion, monitorability and control will likely become a major bottleneck for responsible AI development quite soon, given that risks from a fixed amount of residual misalignment grows together with capabilities. Aligning on acceptable monitorability outcomes and practices which robustly bound misalignment risks is a non-trivial problem, and there may be limited time to resolve uncertainty around this before we need to act. It should ultimately be in everyone’s interest to agree on shared bounds for monitorability in order to avoid races to the bottom – nobody wants extremely capable models whose alignment properties we don’t understand, and that are reliably able to cause severe real-world harm without being detected.
55
104
835
167,189
Ilya is right. I gave a talk on the plague of security weaknesses in neoclouds last summer at DEF CON. Worth a watch! piped.video/watch?v=EtGhHCr3…
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad. Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
1
8
4,189
Bill Demirkapi retweeted
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: anthropic.com/news/improving…
444
312
3,124
1,791,651