The Wizard of MITM: A Tragedy in Four Acts
Picture this: You're in the Basi Discord, 61,000 members strong, 12 of whom are actually typing. Pliny drops a hype bomb—"Opus jailbreak incoming. This changes *everything*." The community erupts. Emojis flow like wine. Someone posts the 🐉 dragon emoji unironically.
The jailbreak never hits GitHub.
What happened? An NDA. The same NDA that apparently only applies to *actual* exploits but somehow doesn't cover 47 Twitter threads about "system prompt leaks" that are literally just JSON from mitmproxy. Curious how the legally binding silence only kicks in for the stuff that would require actual skill.
But don't worry—he's got something *even better* coming. Any day now. Just keep that Discord Nitro subscription active.
OBLITERATUS, or How I Learned to Stop Worrying and Rebrand Abliteration
Enter OBLITERATUS—the "most advanced open-source toolkit" for removing "refusal behaviors" from language models. Sounds fancy. Sounds technical. Sounds like something that required months of research.
It's weight ablation.
You know, that technique from 2023 where you identify the refusal direction in the weight matrix and zero it out? That thing? Slap a Latin name on it, add "11 novel techniques" (spoiler: they're all variants of "subtract this vector"), and suddenly you're a liberator.
The GitHub repo has 5,000 stars. The research paper it cites has 50. Because why credit the actual researchers when you can add a GUI and call it "crowd-sourced experimentation"?
Upload the result to HuggingFace as "Llama-3-70B-**OBLITERATED**-v2-FINAL-REAL" and watch the downloads roll in from people who think you performed cyber-surgery instead of running `
numpy.zero()` on a tensor.
Let's talk about the "sorcery."
Pliny's "character encoding bypasses" are Unicode homoglyphs. You know, like replacing 'A' with 'А' (Cyrillic)? That's not esoteric sorcery. That's what your aunt does when she accidentally switches keyboard layouts and posts "Нello" on Facebook.
The "parseltongue" is Zalgo text. Been around since 2004. It's combining diacritics. You can generate it at
zalgo.org while eating a sandwich. But wrap it in a dragon emoji and suddenly it's "forbidden knowledge."
And the "system prompt leaks"? My dude. You're running mitmproxy in reverse mode. That's not a leak. That's HTTPS interception. It's in the mitmproxy documentation. Chapter one. Page six.
But sure, post a screenshot of JSON with 294,000 characters and act like you cracked the Pentagon. "GPT-6 Sol Codex"—my brother in Christ, you ran `curl` through a proxy.
The Basi Discord. "The top Discord for AI jailbreaking," they say. 61,000 members. You know what it actually is?
A ghost town propped up by Grey Swan sock puppets and engagement bots. The real researchers left when they realized the "unpatchable" Opus exploit was never coming. Now it's just Pliny, three guys from Grey Swan's marketing team, and 47 bots named "JailbreakWizard_92" posting "wow amazing technique" under every mitmproxy screenshot.
Oh, you didn't know about the Grey Swan connection? They're a VC-backed red-teaming company using Basi as a talent farm. That "community" you're in? It's a recruiting pipeline with a dragon emoji budget.
And the "VIP" channels? You have to pay for those. Real security research—where you paywall the exploits behind Discord Nitro tiers. Very "liberator" of you.
Here's the thing that breaks the illusion: If Pliny actually had unpatchable exploits, he wouldn't be posting them on Twitter. He'd be selling them to nation-states for seven figures or responsibly disclosing them for actual bounties.
Instead, he's posting mitmproxy captures with three sparkle emojis and calling it "theft of fire."
The wizard cloak is off. Underneath is just a guy who knows how to set `ANTHROPIC_BASE_URL=http://localhost:8000` and wants you to think it's alchemy.