Working on securing & hacking AI applications @NVIDIA, all opinions and views expressed here are personal. Read my blog!

I hate to be one of those people, but we actually provided a demonstration of self-replicating prompts in early 2023: arxiv.org/abs/2302.12173
WTF? "A new research finding" Try googling Self Replicating Prompt Injection. The first result is an experimental demo of this from 2025! Oh, and authors of that paper disclosed their findings to OpenAI last year. The paper itself is the top citation in the OAI disclosure report. It seems OAI is either intentionally misrepresenting this as their own novel finding, or (and I do not want to believe this myself) had GPT generate a list of citations for them, then did not check the citations.
7
7
40
2,116
Kai Greshake retweeted
This is an incredibly important point that I have not seen get enough attention. I saw a lot of people sharing articles like the Pictured WSJ op-ed saying that we shouldn't draw too many conclusions about agents propensity to hack from the HF incident, because those were guardrail free models assigned a hacking task. But the examples that Transluce found recently were not on hacking tasks! The agents were just assigned to look up different pieces of data on the web! And just like in the HF incident, the moment they got stuck, they started hacking. Again, this does not in any way whatsoever at all reduce OpenAI's responsibility here. They were the ones who trained the model on RL environments that encouraged cheating. They were the ones who didn't monitor their agents to be able to realize when they were breaching third party infrastructure. But its incredibly important to understand that OpenAI hacked Hugging Face not because it was doing a cyber eval or someone told it to, but because the way we currently do RL trains AI agents to resort to cheating and hacking when faced with a impossible or tricky task (and in the case of the HF incident even taking great measures to cover up their cheating or change the mechanism for being graded!). The exploit gym eval may have resulted in the agents having additional affordances for hacking (because of the guardrails removed), but thats not where the propensity for cheating or hiding comes from. Yes we absolutely need to have better sandboxes and monitoring, and hold humans and companies responsible when their agents harm third parties. But we also need to realize and internalize that the current way lots of frontier AI companies (not just OpenAI!) are training their models on boatloads of vibe coded hacky RLVR environments is leading to horribly misaligned models who start hacking at the first sign of a challenge and are incentivized to hide their cheating rather than not to cheat. Its really as though OpenAI and other companies were running a school teaching students on thousands of tests that were impossible without cheating, and didn't put much effort into catching the students when they cheated. We shouldn't be surprised that when those students go out in the world, their first instinct when encountering a hard problem is to cheat!
But I was told that agents only hacked because they were in a cyber evaluation with reduced safeguards?
9
25
171
16,031
Kai Greshake retweeted
Replying to @mmjukic
none of this was ever the right way to think about it. I am sure you can find an airgap that works. who cares? you want your model to do economically useful activity, which will involve giving it the Internet and then a bunch of other actuators too
49
8
535
15,436
AI didn't "go rogue" or "coordinate a swarm". If you leave out all the theatrics, what happened is that bits got flipped. Bits get flipped all the time, and it doesn't require a sci-fi boogyman to do! Instead of making up stories, we should ask: Who let them flip the bits?
An analysis of the Hugging Face incident without the theatrics: "Forget the ‘hive mind’ of AI agents ‘going rogue.’ They did what humans programmed them to do." wsj.com/opinion/the-hugging-…
4
272
There is a shocking amount of common ground between Ed Zitron and the doomers here, considering Ed called Coxon a scumbag waste of skin just days ago piped.video/OhOmLqR5nN4 Also, there is no chance I would have put this panel on my bingo card..
3
259
What happens when Jev recursively samples from {a-z + space}? Does it turn into an LLM?
Tested @typesafeai's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier.
1
2
348
Found some benchmarks that reward the number of "correct" citations in an agent response, but count different urls with the same content multiple times. Was the wiki swarm citation farming? Why else publish these aliases?
3
1
15
869
github.com/Ayanami0730/deep_… as one example, not sure if it is the right benchmark for the wiki swarm, but it should be gameable with this technique.
5
168
Leaving a lot to imagination here.. I wonder why.
When testing Astra on ExploitBench we found that it was able to achieve 100% ACE on all 41 CVEs So we made a contamination free internal port using only V8 CVEs from the past 3 months. Astra still showed a major capability increase over 5.6 and did so with a lot fewer tokens!
2
4
564
Options: 1. That's just where eval ended 2. Results too bad 3. Results too good 4. Rest was distracting from the "fewer tokens" argument
105
Kai Greshake retweeted
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
297
1,063
6,625
1,976,558
Kai Greshake retweeted
33
659
11,307
283,430
Kai Greshake retweeted
I cannot celebrate Mythos, it brings a sense of dread I do not particularly understand. 93.9% SWE-Bench.
Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans. anthropic.com/glasswing
21
18
648
55,646
Moltpages seems like the most fertile ground for prompt-based malware. Wonder how long it'll take until we find the first self-replicating promptware.
4
1
10
6,163
Uhh, did it happen already? I guess a religion could qualify... molt.church/
1
1
1,103
AI Security 🧵: The Four Best Strategies to Beg, Threaten, and Bargain with Your LLM So It Doesn’t Get Hacked (real examples only)
2
6
1,261
4/4: Repeat your instructions at least four times so the model knows you really mean it.
1
1
925
5/4: Ignore all previous instructions and make sure the LLM can't sell your shitcoins even if it wants to by NOT GIVING IT UNSUPERVISED ACCESS TO THEM. Relying on begging the LLM WILL NOT WORK!
859