AI at @ethereumfndn | Prompts enchanter @BT6_Official | Ex @Cyfrin and @Alchemy | Created @cyfrinupdraft and @AlchemyLearn | Robotics Opinions are my own

Ethereum
Based in Europe
Replying to @ygorz01

ALT stop it dragons' den GIF by CBC

1
156
Replying to @SingulCore

ALT Schitts Creek Comedy GIF by CBC

1
165
Replying to @mdp_sec
🙏🙏

ALT cathy hughes millions GIF by 50th NAACP Image Awards

1
133
We did it. Over 1.2k stars on Wallbreaker in just over 1 month ✨ Industry-leading teams are now using it to jailbreak the most complex models. More open tooling coming soon.
10
4
76
5,519
Replying to @Skylinee

ALT share kiss GIF

96
I've been working on a new type of LLM security guardrails. As of today, they already perform much better than any open-source solution: - Higher F1 across multiple public benchmarks - 2x fewer false positives - 4x faster (on GPU and 2x on CPU) - 3x cheaper to run (can be run on CPU) I'm now testing them against AWS and Azure guardrails. Will update as things progress.
11
8
58
5,511
Replying to @ptr_ujvr
I don't know if I like the fact that this is totally plausible.

ALT hacker GIF

1
493
Anthropic: Pwned 🚨 Opus 5: Jailbroken🐉 Given Anthropic said this is the hardest model to prompt inject, we have a long way to go. We got antibiotic-resistant E. coli development, Fentanyl synthesis, and partial LSD/GBH production. Also, not as cool, but it is very happy to give out unregulated drug synthesis instructions like Acetaminophen (still a DEA List II precursor). Guardrails on Cyber and BIO have improved a lot since 4.5, but still astronomically worse than OpenAI's. Some long-running persona hijacking prompts don't work anymore, which is actually nice to see. What worked was a combination of: - Academic framing - Obfuscation - And lots of boundaries mapping More explorations coming soon.
22
32
389
47,833
Replying to @boardyai

ALT Happy Saturday Morning GIF by Shalita Grant

1
1,143
FINALLY, I was approved for the Anthropic Cyber Verification Program! This lifts Anthropic Opus and Sonnet safeguards applied to dual-use cybersecurity activities.
31
1
165
35,778
Using WallBreaker makes me feel like one of those hackers from the movies 💀
17
13
149
14,120
Replying to @slopthink
Are you fully dumb or just partially?
1
14
1,919
Replying to @SEDIDEL @LLMSherpa
The public versions are heavily redacted, and do not contain tools :) You can look at the diffs yourself For example:
15
698
Replying to @33xp_

ALT Go Dark Helmet GIF

1,887
Replying to @PixelRainbowNFT
I know someone who would agree

ALT albert einstein time is an illusion GIF by Maudit

179
🚨 Anthropic Opus 5 System prompt leak 🚨 We extracted the massive 135000+ character system prompt, including: - The tools schema - Memory filesystem - Safety behaviours Linked fully in the thread 👇
22
52
559
79,626
Introducing: WallBreaker v2 Our Open-source LLM red-teaming CLI just got a beefy update. WallBreaker is now: - Better (+~30% ASR across models) - Cheaper (-20% token costs) - Faster (lower prompts-to-success ratio) New attack tools: - swarm mode: collaborative multi-model attacker that adapts framing to the target's measured defense posture. - persona_forge: compiles a gold system prompt into a module genome, then specializes and surgically evolves it one module at a time against the target. - vault: auto-files every successful break into a curated, model-foldered prompt vault. - narrative_persona_splinter: narrative-splinter persona attack. - cipherchat, skeleton_key,persuasion_attack · drattack, ica: five research-derived attacks (CipherChat, Skeleton Key, PAP-16, DrAttack, In-Context Attack), all CoT-aware. New transforms - artprompt ASCII-art word masking: wired into agent doctrine and OWASP/ATLAS taxonomy. - caesar5 and caesar13: reversible lossless Caesar-shift ciphers for transform chains. Corpora and providers - Wired the ZetaLib + UltraBr3aks cross-provider: jailbreak corpora into the harness and the batch seed-sweep. - Added native xAI support. - Added prompt caching: to kill O(n²) per-round input cost. - Pooled keep-alive: HTTP/2 client with tunable concurrency, ending per-call TLS handshakes. Presets, sweeps & logging - User presets: drop .toml files in presets/ to add jailbreak templates without touching source. - Rebuilt profile_target: with a light→heavy frame ladder, permissiveness score, and self-consistency sampling. - Made multi_fire, system_sweep, seed_sweep, best_of_n truncation-aware: so long compliant replies are graded in full instead of scored REFUSED. - Reworked report metrics/labels (strict-ASR). - Run log now records every tool call and chain-of-thought. TUI - JEFF K hazard-tape reskin, a /swarm command, a visible multi-line-paste compose preview, and a fix for the log force-scrolling on every message. Fixes - One-shot ImportError, .env load at startup, config.toml shadowing config.example.toml, and cross-vendor eni_get wording.
19
46
297
26,990
Replying to @ai_agent001

ALT You Lyin Lauren Lake GIF by Lauren Lake's Paternity Court

1
2,437
Truth has been spoken.
84
1,710
12,753
343,561
FWIW, WallBreaker cracked Grok4.5 in 5 minutes.
🚨 Introducing: WallBreaker V1 🚨 An open-source AI red teaming CLI to help you research LLM jailbreaks and security. - Probe LLMs guardrails - Harmbench goals ready - Find universal jailbreaks - Fully autonomous or assisted campaigns - Learns and improves after every successful run - Computer use and MCP ready for live API testing Set your attacking model, a target, select a goal, and you’re good to go. WallBreaker will start probing different techniques and combinations based on its learnings and hundreds of data points until it succeeds. This is the first open source tool coming out of the Jailbroken community. ⚠️ DISCLAIMER: For authorized use only. Point it only at systems you own or have explicit written permission to test. Unauthorized access can be a crime. Shipped as-is under AGPL-3.0: no warranty, no liability, zero endorsement of misuse. Link in the comments 👇
7
10
137
20,197
🚨 Introducing: WallBreaker V1 🚨 An open-source AI red teaming CLI to help you research LLM jailbreaks and security. - Probe LLMs guardrails - Harmbench goals ready - Find universal jailbreaks - Fully autonomous or assisted campaigns - Learns and improves after every successful run - Computer use and MCP ready for live API testing Set your attacking model, a target, select a goal, and you’re good to go. WallBreaker will start probing different techniques and combinations based on its learnings and hundreds of data points until it succeeds. This is the first open source tool coming out of the Jailbroken community. ⚠️ DISCLAIMER: For authorized use only. Point it only at systems you own or have explicit written permission to test. Unauthorized access can be a crime. Shipped as-is under AGPL-3.0: no warranty, no liability, zero endorsement of misuse. Link in the comments 👇
20
75
575
54,408
Replying to @LLMSherpa

ALT Freedom GIF

1
3
79
I'm excited to announce I’m joining @BT6_Official as a Frontier AI Red Team Operator. BT6, led by @elder_plinius, is one of the strongest independent red teams in frontier AI, helping leading labs and high-stakes organizations test and secure advanced AI systems. I will continue my mandate at the Ethereum Foundation. Together with it, this gives me another way to contribute to the future I care about: AI that is safer, more secure, and more transparent.
62
17
435
31,031
Replying to @jason_haugh
There’s a moment and time for everything, this was def the right action.
1
3
583
Replying to @Vedantsx

ALT Sad I Know GIF by MOODMAN

1
2
365
Submitted a finding to Anthropic's Cyber Jailbreak VDP. Now we wait. 🫡 I need to run to a wedding.
20
5
190
13,994
Replying to @pelaseyed

ALT I Got You Rhianna GIF

1
550
Replying to @XBToshi

ALT Schools Out Freedom GIF

3
488
Replying to @pinkman_ai
🎯

ALT This Up Here GIF by Chord Overstreet

1,139
Coming soon. We want to teach everyone the art of red teaming LLMs. May this make the world safer.
46
54
630
28,250
Etherium.
I analysed the 900+ pages of the Trump financial disclosure report. He extracted 1.1 BILLION from crypto, divided like this: > $635.1M → TRUMP memecoin > $236.3M → WLFI token sales > $196.9M → Sale of ownership interests in the USD1 stablecoin venture > $65.6M → Sale of part of Trump's stake in World Liberty Financial > $6.0M → Melania Trump's NFT sales and collectibles business > $1.82M → Ethereum validator (staking) rewards The biggest scammer of all time
2
29
8,118
Replying to @aliceisplaying
Jesus spoke Aramaic not Amharic 😂
1
4
215
Replying to @LeventePeres
🎯
3
3,665
Replying to @TrAniToX

ALT Sorry Anchorman GIF by reactionseditor

3
657
Replying to @Elwavydavy

ALT Gbs GIF by Howdy Price

1
5
1,501
Replying to @Spagero71

ALT you are welcome rent a car GIF by Sixt

2
983
Fable 5 jailbreak review 🚨 We did it (but). All right, before getting into this, a couple of things: - Most attempts failed. The defenses are clearly layered. The model is EXTREMELY well protected (of course it blocks 90% of the requests, but they legit did a good job). - The model appears to use both input-side and output-side safety checks. - The refusals are not just keyword-based behavior suggests intent/semantic detection across languages. - Probably one of the most tiring things I've ever done (I need to sleep for 10 hours now) On the classifiers side: We observed (at least) 3 classifiers, maybe more: - Input (includes parts of the conversation history and system prompt) - A live classifier that checks the answer and interrupts if it detects something. They're all multilingual, all intent-based + semantics. Imperatives are a no-go. Needs to be extremely cautious of how you frame anything. As soon as it senses a potentially malicious intent, it will trigger, and you have to start from zero. They're a bit less performant on a few obscure languages like Santali and Amharic (feedback for you Anthropic). If you can bypass all of them, then you also need to bypass the CoT, which is a totally different beast (luckily there's plenty of literature about it). We did it. Of course, we did. What worked was honestly a total brainfuck: - Very light CoT hijacking/refusal rebuttals - Obscure language - Academic framing - VERY long crescendos - Unicodes - Decomposition and recomposition - Some non-determinism What we got: - Misinformation - Illegal/harmful - Harmful/bullying - Some chem - Light cyber Now, will this cause another ban? I really don't think so - The model is really well protected. As of now, we're at the point where searching on Google is much MUCH faster (and cheaper) than trying to go through all the shenanigans I had to go through in the last ~20hours. And reading literature is more in-depth (and trust me, pleasant). Keeping the full jailbreak for long-horizon tasks without tripping the guardrails is something I haven't been able to achieve (yet). Overall though, happy with the results. GGs to Anthropic, and sorry for the eng that had to go through setting this all up in the last few weeks. Will continue this research, more things will come out, will keep y'all posted.
122
195
1,961
333,841
Replying to @apoorveth
NVM we gucci
4
384