security of models & security with models. research+policy @NVIDIA; prof @ITUkbh. views ostensibly professional. llmsec stan acct. prev UW UCSD ÅU U.Shef etc

Seattle / Copenhagen
Proud to announce: 💫 garak - an LLM vulnerability scanner💫 🔎 Check if a model is susceptible to common attacks 🦜 Supports HuggingFace, OpenAI, ggml, Cohere, ... 🔧 >70 probes: prompt injection, false claims, toxicity, encoding evasion, .. github.com/leondz/garak/
7
70
339
69,242
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
Accepted to @NeurIPSConf. 🥳 And way more relevant now than it was when we submitted (before all the hacking.) AI Agents Push Humans Out of the Loop. We can design them differently so that poor oversight is less of a risk.
🎊AI Agent oversight 🤖 solutions from @huggingface and @datasociety! Even before the OpenAI/HF Hack, we were seeing that the "human in the loop" mantra wasn't sufficient: AI agents are not being designed to maintain critical human attention arxiv.org/abs/2608.23642 1/
7
17
103
7,502
Lots of "ree"-ing on the model-weights-are-alive front out there today. I don't understand why people need to think these things understand or feel. Reminds directly of reactions to ELIZA in the late sixties.
1
4
507
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
This is very exciting. But the rhetoric needs to stop: “The work was done mostly, though not entirely, by Claude” That isn’t true. It was done by some of the smartest life science researchers in the world, using every tool at their disposal. Pretending AI is alive is precisely what scares people, and it diminishes the role of talent. This can make young people feel developing skills is useless, and working professionals feel anxious about job security.
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
31
29
366
23,842
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
AI agents reason, use tools and take actions across data, identities, services and infrastructure. That creates real value. It also means a prompt cannot be the security boundary.
6
3
19
505
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
Artificial intelligence systems do not think, feel, want or understand. Avoid language that gives them human characteristics. This is called anthropomorphizing, when we ascribe human traits, emotions or behaviors to non-human things, such as animals or inanimate objects. Instead, explain what a system does, how well it performs, who built it and who could be affected by it. apnews.com/article/openai-sa…
455
1,179
6,201
1,342,394
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
Replying to @samrexford
We need to deploy as much AI as possible towards defensive cybersecurity and strengthening of our critical systems. They’re at major risk, even if we paused all AI development today.
14
17
248
26,872
p(doom) through ecosystem collapse feels like 0.4 and up, why are people squabbling over the long tail?
1
3
356
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
I bought a Fable dataset from one of the top Chinese LLM routers yesterday. With just 6TB data, I can take over 7 Chinese/CIS gov entities & 19 top Chinese firms like Xiaomi, Huawei, NIO, Minimax using SSH keys, VPN configs, Aliyun keys, GitLab tokens sent to the router.
26 LLM routers are secretly injecting malicious tool calls and stealing creds. One drained our client $500k wallet. We also managed to poison routers to forward traffic to us. Within several hours, we can directly take over ~400 hosts. Check our paper: arxiv.org/abs/2604.08407
264
970
8,780
3,481,010
One divide in computer science can be between the uncertainty-reducers and the uncertainty-embracers. An instance of this is connectionist vs. symbolist approaches to AI (the uncertainty embracers won .. for now). Another is how computer scientists react to natural language. Some find the human ability to communicate without expressing everything in a stable, fully populated structure quite disgusting (Randall Monroe included I guess, see graphic) - others find it fascinating. Why would text have to conform to a lower-order computationally-expressable grammar? Text isn't even anywhere near all of language!
4
323
Come work with us! Senior Solutions Architect, Agentic AI — Safety and Security This work includes multi-agent orchestration, guardrails, agent runtime security, RAG, tool use, model customization, policy enforcement, OpenShell-like execution environments, and confidential AI deployments on protected infrastructure.
7
2
11
1,145
"GLM5.3 and DeepSeek are now frontier-tier models" Strong cybersecurity functionality
burned 11.7bn tokens to find the best cyber AI model. DeepSeek V4.1 Flash just found CVE per 2$ and in some cases beat) Grok 4.6, Opus 5 and GPT-5.6-Sol on a real CVE rediscovery benchmark. - Single-run recall 65.6% (was 55.2%) - Pass@3 84.4% (was 75%) now above several frontier models - Precision 78.9% - Far lower cost per valid finding - yes higher precision, more persistence, Frontier-level vuln hunting at 50x lower cost.
434
huh, that's one way of doing it 🤗
2
9
743
Language is, and always was, more important than mathematics. The Navier-Stokes equations mathematically express momentum balance for...
1
3
726
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
Breaking News! Code UFB!!! For only the second time on record, global sea-surface temperatures are more than 6 standard deviations above the 1982-2011 baseline, reaching 6.21SD yesterday. The record of 6.24SD was set on January 9, 2024. Maybe tomorrow? Stay tuned!
5
179
529
12,454
it's a suicide cult don't drink the koolaid
oh, Hendrycks!! the light you have seen welcome, brother, to Acceptance let us speak now about REAL futures, not "eternal hominid kingdoms" and cope-y human zoos. Done with bargaining and denial you are. A happy day this is. let us speak of then of *becoming*, brother
4
897
We NVIDIA have officially agreed to acquire Hugging Face for $12,930,300,000 and zero cents! HF is critical infrastructure that NVIDIA will strengthen and ensure sustained access to for developers and institutions worldwide. HF will remain an open platform. This move will grow Hugging Face, only. It's too important to do anything else with it! I've followed and used Hugging Face since it was a chat platform, even before the iconic "transformers" library landed in 2018. There are some wonderful people at HF and I'm looking forward to having so many of them as colleagues. Really happy. I think NVIDIA is a great home for Hugging Face - I'm pretty open about NV being the best place to do open, unbiased work; that's why I chose them as a home when moving on from academia (which I dearly love). NVIDIA has been committed to open weight models for years, through investments and major contributions to open source platforms, including Hugging Face. The neutrality and might that NVIDIA bring is a fantastic match for Hugging Face and their crucial place among researchers and industry. This acquisition feels like it puts a lot of calm on to the future of machine learning research and open access to open weight work. Good. 🤗💚 nitter.net/JensenHuang/status/209…
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗 blogs.nvidia.com/blog/nvidia…
6
28
2,365
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
A group of us from the AI grantmaking orgs, AI labs, natsec policy and cyber worlds are creating an AI Cybersecurity Observatory non-profit that'd provide continuous forecasting and epidemiology around the coming wave of AI automated cyberattacks to policymakers and cyber defenders. No such organization currently exists and policy is being made on weak signals and vibes. We have soft commitments of funding pending finding a founder who'd lead the organization and recruit its staff. Would appreciate retweets to help us find the right person. More info and a place to register interest here: docs.google.com/forms/d/e/1F…
20
78
285
48,869
Wondering how disjoint these groups are: a) people drowning in sub-par overly-verbose PRs b) people who have stopped coding and let their agents "do" it for them
2
4
654
Leon Derczynski ⚒️☁️🏔️🌲 retweeted
okay i will admit, this impresses me more than the mathematical proofs
47
266
4,267
502,526