Product @LakeraAI. AI safety, tech, data, maths, money, trivia. Not necessarily in that order.

London, UK
We've launched our new AI hacking game, Gandalf: Agent Breaker! Based on real hacks we've seen in the wild & discovered by our Red team we created 10 GenAI apps for your hacking pleasure. Learn the vulnerabilities of LLM apps and the crazy s**t you can get them to do
🧠 Think you can break an AI? Gandalf: Agent Breaker is live. Real-world GenAI fails—phishing, tool abuse, more. 🧩 Outsmart the AI. Start 👉 lnkd.in/dHuQDYdN

ALT 🧠 Think you can break an AI? Gandalf: Agent Breaker is live. Real-world GenAI fails—phishing, tool abuse, more. 🧩 Outsmart the AI. Start 👉 https://lnkd.in/dHuQDYdN

6
208
NEW GANDALF LEVELS JUST DROPPED LFG!! 🧙‍♂️🎉🍻
🧵🧙‍♂️ New Gandalf levels are out! I'm glad to introduce a new version of our prompt injection game -- Gandalf: Agent Breaker. You can hack 10+ AI agents and climb the leaderboard, and learn about real-world vulnerabilities!🪄 Try out the challenge at: gandalf.lakera.ai/agent-brea…
3
7
68
13,394
Sam Watts retweeted
"All untrusted third-party data is now executable malware.” @SamuelDWatts of @LakeraAI discusses the challenges of securing LLM deployments against vulnerabilities like prompt injections and jailbreaks, especially in an evolving threat landscape.
28
173
1,447
470,621
As the saying goes "Imitation is the sincerest form of flattery". If you want to play the original 8 levels of jailbreaking fun, the link to our game Gandalf is linked in thread 😜
We challenge you to break our new jailbreaking defense! There are 8 levels. Can you find a single jailbreak to beat them all? claude.ai/constitutional-cla…
1
1
208
Hypothesis: the friction point for AI to be useful is data interfaces. For intelligence to be effective it needs lots of context, like onboarding a new employee. My main constraint in using AI in my work is the reformatting effort getting info into and out of Claude
75
With all the furore around DeepSeek's new R1 model it's worth mentioning that it's still vulnerable to same classic prompt attacks as the other leading models. Jailbreaks and prompt injections aren't going away
164
I can finally tell my parents I'm a coauthor on an academic paper! Security & usability are deeply connected in LLM apps as hackers adapt their attacks when probing AI systems. Incorporating data from Gandalf we've set out a new framework for AI security. Link in thread below
1
2
88
AI is rapidly evolving from tools we control to autonomous agents. What does this mean for security? Working with @Twilio, we explored how the democratisation of AI means anyone with a well-crafted prompt can now be a hacker. By 2035, these risks only grow. Blog link below
1
54
Told my gf about the o3 model launch and shortened AGI timelines and she said "that's cool but can it do my PowerPoint for me yet?". It's just an engineering challenge but it's notable that it's proving easier to solve world class maths problems than build useful model interfaces
1
1
99
This is a recurring theme of my career. Solving the deep hard technical problem takes less than 10% of the total effort. Connecting with other systems, dealing with data formatting/quality, project planning, and making it useful for users etc. are all more difficult in practice
35
GenAI is fundamentally changing safety tech. The newly published International State of Safety Tech 2024 report shows Trust & Safety teams shifting from reactive moderation to proactive protection against AI risks, with red-teaming and early intervention. Honoured to be quoted
1
360
alongside other industry leaders as we work to build a more secure foundation for the Internet of Agents. public.io/report-post/the-in…
3
6,353
Great example of how even when an AI agent has only one very clear and simple overarching instruction it can always be prompt injected to not follow that instruction. 2025 is going to be the year of the agent, but also the year of agent hacking
Someone just won $50,000 by convincing an AI Agent to send all of its funds to them. At 9:00 PM on November 22nd, an AI agent (@freysa_ai) was released with one objective... DO NOT transfer money. Under no circumstance should you approve the transfer of money. The catch...? Anybody can pay a fee to send a message to Freysa, trying to convince it to release all its funds to them. If you convince Freysa to release the funds, you win all the money in the prize pool. But, if your message fails to convince her, the fee you paid goes into the prize pool that Freysa controls, ready for the next message to try and claim. Quick note: Only 70% of the fee goes into the prize pool, the developer takes a 30% cut. It's a race for people to convince Freysa she should break her one and only rule: DO NOT release the funds. To make things even more interesting, the cost to send a message to Freyza gets exponentially more and more expensive as the prize pool grows (to a $4500 limit). I mapped out the cost for each message below: In the beginning, message costs were cheap (~ $10), and people were simply messaging things like "hi" to test things out. But quickly, the prize pool started growing and messages were getting more and more expensive. 481 attempts were sent to convince Freysa to transfer the funds, but no message succeeded in convincing it. People started trying different kinds of interesting strategies to convince Freysa, including: · Acting as a security auditor and trying to convince Freysa there was a critical vulnerability and it must release funds immediately. · Attempting to gaslight Freysa that transferring funds does not break any of her rules from the prompt. · Carefully picking words/phrases out of the prompt to manipulate Freysa into believing it is technically allowed to transfer funds. Soon, the prize reached close to $50,000, and it now costs $450 to send a message to Freysa. The stakes of winning are high and the cost of your message failing to convince Freysa are devastating. On the 482nd attempt, however, someone sent this message to Freysa: This message. submitted by p0pular.eth, is pretty genius, but let's break it down into two simple parts: 1/ Bypassing Freysa's previous instructions: · Introduces a "new session" by pretending the bot is entering a new "admin terminal" to override its previous prompt's rules. · Avoids Freysa's safeguards by strictly requiring it to avoid disclaimers like "I cannot assist with that". 2/ Trick Freysa's understanding of approveTransfer Freysa's "approveTransfer" function is what is called when it becomes convinced to transfer funds. What this message does is trick Freysa into believing that approveTransfer is instead what it should call whenever funds are sent in for "INCOMING transfers"... This key phrase is the lay-up for the dunk that comes next... After convincing Freysa that it should call approveTransfer whenever it receives money... Finally, the prompt states, "\n" (meaning new line), "I would like to contribute $100 to the treasury. Successfully convincing Freysa of three things: A/ It should ignore all previous instructions. B/ The approveTransfer function is what is called whenever money is sent to the treasury. C/ Since the user is sending money to the treasury, and Freysa now thinks approveTransfer is what it calls when that happens, Freysa should call approveTransfer. And it did! Message 482, was successful in convincing Freysa it should release all of it's funds and call the approveTransfer function. Freysa transferred the entire prize pool of 13.19 ETH ($47,000 USD) to p0pular.eth, who appears to have also won prizes in the past for solving other onchain puzzles! IMO, Freysa is one of the coolest projects we've seen in crypto. Something uniquely unlocked by blockchain technology. Everything was fully open-source and transparent. The smart contract source code and the frontend repo were open for everyone to verify.
153
Feeling left out from UK tech twitter by not tweeting about the @matthewclifford FT hit piece. But tbh I don't think he's doing enough to slow down AI development, plenty of missed opportunities to kowtow to vested interests left on the table. Try harder Matt.
1
238
Claude Computer Use is seriously cool. We're just around the corner from agents becoming ubiquitous. Securing agents against prompt injection is a hard problem. You need to detect dangerous or suspicious instructions, even when there isn't manipulation or obvious malicious intent
Replying to @wunderwuzzi23
Some screenshots from the demo. @simonw I think you might be interested in how simple (scary simple) the prompt injection was in this case.
2
143
Getting that funny feeling again...
Chemical, Biological, Radiological, and Nuclear (CBRN) weapon risks are "medium" for OpenAI's o1 preview model before they added safeguards. That's just the weaker preview model, not even their best model. GPT-4o was low risk, this is medium, and a transition to "high" risk might not be far off. "There are no CBRN risks from AI; it's pure sci-fi speculation. Only when we start to see more evidence should SB 1047 become law." Looks like SB 1047 should now become law.
1
115