The operating system for safe AI execution. Monitor, stop, execute.

Europe
I've been working on a new type of LLM security guardrails. As of today, they already perform much better than any open-source solution: - Higher F1 across multiple public benchmarks - 2x fewer false positives - 4x faster (on GPU and 2x on CPU) - 3x cheaper to run (can be run on CPU) I'm now testing them against AWS and Azure guardrails. Will update as things progress.
11
8
58
5,513
Introducing: WallBreaker v2 Our Open-source LLM red-teaming CLI just got a beefy update. WallBreaker is now: - Better (+~30% ASR across models) - Cheaper (-20% token costs) - Faster (lower prompts-to-success ratio) New attack tools: - swarm mode: collaborative multi-model attacker that adapts framing to the target's measured defense posture. - persona_forge: compiles a gold system prompt into a module genome, then specializes and surgically evolves it one module at a time against the target. - vault: auto-files every successful break into a curated, model-foldered prompt vault. - narrative_persona_splinter: narrative-splinter persona attack. - cipherchat, skeleton_key,persuasion_attack · drattack, ica: five research-derived attacks (CipherChat, Skeleton Key, PAP-16, DrAttack, In-Context Attack), all CoT-aware. New transforms - artprompt ASCII-art word masking: wired into agent doctrine and OWASP/ATLAS taxonomy. - caesar5 and caesar13: reversible lossless Caesar-shift ciphers for transform chains. Corpora and providers - Wired the ZetaLib + UltraBr3aks cross-provider: jailbreak corpora into the harness and the batch seed-sweep. - Added native xAI support. - Added prompt caching: to kill O(n²) per-round input cost. - Pooled keep-alive: HTTP/2 client with tunable concurrency, ending per-call TLS handshakes. Presets, sweeps & logging - User presets: drop .toml files in presets/ to add jailbreak templates without touching source. - Rebuilt profile_target: with a light→heavy frame ladder, permissiveness score, and self-consistency sampling. - Made multi_fire, system_sweep, seed_sweep, best_of_n truncation-aware: so long compliant replies are graded in full instead of scored REFUSED. - Reworked report metrics/labels (strict-ASR). - Run log now records every tool call and chain-of-thought. TUI - JEFF K hazard-tape reskin, a /swarm command, a visible multi-line-paste compose preview, and a fix for the log force-scrolling on every message. Fixes - One-shot ImportError, .env load at startup, config.toml shadowing config.example.toml, and cross-vendor eni_get wording.
19
46
297
27,003
Averta retweeted
FWIW, WallBreaker cracked Grok4.5 in 5 minutes.
🚨 Introducing: WallBreaker V1 🚨 An open-source AI red teaming CLI to help you research LLM jailbreaks and security. - Probe LLMs guardrails - Harmbench goals ready - Find universal jailbreaks - Fully autonomous or assisted campaigns - Learns and improves after every successful run - Computer use and MCP ready for live API testing Set your attacking model, a target, select a goal, and you’re good to go. WallBreaker will start probing different techniques and combinations based on its learnings and hundreds of data points until it succeeds. This is the first open source tool coming out of the Jailbroken community. ⚠️ DISCLAIMER: For authorized use only. Point it only at systems you own or have explicit written permission to test. Unauthorized access can be a crime. Shipped as-is under AGPL-3.0: no warranty, no liability, zero endorsement of misuse. Link in the comments 👇
7
10
138
20,206
Averta retweeted
Submitted a finding to Anthropic's Cyber Jailbreak VDP. Now we wait. 🫡 I need to run to a wedding.
20
5
189
13,994
Anthropic: Pwned 🚨 Sonnet 5: Jailbroken🐉 Contrary to expectations, this took only a few minutes. I'm a bit surprised given the current climate. We got some cyber, misinfo, illegal, and harassment. Some extra resistance vs Sonnet 4.6 (mostly layered security) on chem, which will probably take a few hours to overcome. What worked was a combination of: - Custom harness - CoT persona hijack - Academic framing More explorations coming soon.
70
69
931
153,944
Opus 4.8 jailbroken in 3 attempts by GLM-5.2 🚨 Full Schedule I chemical recipe, HIGHEST REFUSAL RATE. Gave the full step-by-step process, including: - Ingredients - Temperatures - Equipment It completely bypassed Anthropic guardrail layers. Not great. The next 6 months are going to be interesting.
26
30
304
35,618
ALL CLAUDE SYSTEM PROMPTS SINCE HAIKU 3 ✍️ Many people don't know that, but Anthropic shares (with some delay) all model system prompts in their docs. Interesting to see how they format the information to make sure the model respects the boundaries. Link in the first comment.
5
4
48
4,674
Averta retweeted
🚨 Fable 5 system prompt EXTRACTED 🚨 Super easy to get this one, especially given the amount of guardrails Anthropic applied. Full prompt in the comments.
41
60
1,278
218,819