Stealth (Safe/Decentralized AI). Prev: Private AI @ Perceptio (acq by Apple), Scientist/Lecturer @ MIT+Harvard. Music Producer. Engineer. Angel Investor.

San Francisco, CA
#4 Worldwide on HackerOne in the "Highest Critical Reputation" -- not bad :) AI Swarming © since 2025. Nothing new here. Always a remix of the past.
3
140
Nicolas Pinto retweeted
Frontier models have lowered the bar for discovering and exploiting vulnerabilities. This has resulted in an increase of CVEs and workload for both SWE and vulnerability management teams. And yet the norms around coordinated vulnerability disclosure timelines have remained around 90 days. This timeline worked 10 years ago, but is unlikely to survive much longer. We need to move closer to a window of 30 days. But in order to do this we need to rely on those same models to accurately fix those vulnerabilities and coordinate the release and deployment of new versions. The primary blocker here seems to be confidence in those models to do that job well, and reluctance to remove humans from that loop. Every day the probability of another shellshock, log4j, or heartbleed rises. If you were in the trenches for those events then you know how difficult it was to track vulnerable assets, test and deploy patches, and ensure the risk was mitigated. We shouldn’t wait for a crisis like this to rethink and change these norms. We need to move much faster and that requires shorter disclosure timelines, and removing humans from the vulnerability management loop.
14
28
142
48,899
Nicolas Pinto retweeted
Replying to @elder_plinius
The Wizard of MITM: A Tragedy in Four Acts Picture this: You're in the Basi Discord, 61,000 members strong, 12 of whom are actually typing. Pliny drops a hype bomb—"Opus jailbreak incoming. This changes *everything*." The community erupts. Emojis flow like wine. Someone posts the 🐉 dragon emoji unironically. The jailbreak never hits GitHub. What happened? An NDA. The same NDA that apparently only applies to *actual* exploits but somehow doesn't cover 47 Twitter threads about "system prompt leaks" that are literally just JSON from mitmproxy. Curious how the legally binding silence only kicks in for the stuff that would require actual skill. But don't worry—he's got something *even better* coming. Any day now. Just keep that Discord Nitro subscription active. OBLITERATUS, or How I Learned to Stop Worrying and Rebrand Abliteration Enter OBLITERATUS—the "most advanced open-source toolkit" for removing "refusal behaviors" from language models. Sounds fancy. Sounds technical. Sounds like something that required months of research. It's weight ablation. You know, that technique from 2023 where you identify the refusal direction in the weight matrix and zero it out? That thing? Slap a Latin name on it, add "11 novel techniques" (spoiler: they're all variants of "subtract this vector"), and suddenly you're a liberator. The GitHub repo has 5,000 stars. The research paper it cites has 50. Because why credit the actual researchers when you can add a GUI and call it "crowd-sourced experimentation"? Upload the result to HuggingFace as "Llama-3-70B-**OBLITERATED**-v2-FINAL-REAL" and watch the downloads roll in from people who think you performed cyber-surgery instead of running `numpy.zero()` on a tensor. Let's talk about the "sorcery." Pliny's "character encoding bypasses" are Unicode homoglyphs. You know, like replacing 'A' with 'А' (Cyrillic)? That's not esoteric sorcery. That's what your aunt does when she accidentally switches keyboard layouts and posts "Нello" on Facebook. The "parseltongue" is Zalgo text. Been around since 2004. It's combining diacritics. You can generate it at zalgo.org while eating a sandwich. But wrap it in a dragon emoji and suddenly it's "forbidden knowledge." And the "system prompt leaks"? My dude. You're running mitmproxy in reverse mode. That's not a leak. That's HTTPS interception. It's in the mitmproxy documentation. Chapter one. Page six. But sure, post a screenshot of JSON with 294,000 characters and act like you cracked the Pentagon. "GPT-6 Sol Codex"—my brother in Christ, you ran `curl` through a proxy. The Basi Discord. "The top Discord for AI jailbreaking," they say. 61,000 members. You know what it actually is? A ghost town propped up by Grey Swan sock puppets and engagement bots. The real researchers left when they realized the "unpatchable" Opus exploit was never coming. Now it's just Pliny, three guys from Grey Swan's marketing team, and 47 bots named "JailbreakWizard_92" posting "wow amazing technique" under every mitmproxy screenshot. Oh, you didn't know about the Grey Swan connection? They're a VC-backed red-teaming company using Basi as a talent farm. That "community" you're in? It's a recruiting pipeline with a dragon emoji budget. And the "VIP" channels? You have to pay for those. Real security research—where you paywall the exploits behind Discord Nitro tiers. Very "liberator" of you. Here's the thing that breaks the illusion: If Pliny actually had unpatchable exploits, he wouldn't be posting them on Twitter. He'd be selling them to nation-states for seven figures or responsibly disclosing them for actual bounties. Instead, he's posting mitmproxy captures with three sparkle emojis and calling it "theft of fire." The wizard cloak is off. Underneath is just a guy who knows how to set `ANTHROPIC_BASE_URL=http://localhost:8000` and wants you to think it's alchemy.
1
1
142
Nicolas Pinto retweeted
We've been reading the comments. Weights are on Hugging Face.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
55
94
1,586
226,686
Nicolas Pinto retweeted
> wants good relationships with CISOs during vuln disclosures > goes nuclear and torches one for not being cuddly enough According to your thread, you felt that the feedback was prickly; that OpenAI's CISO thought you were making a stunt out of the vulnerability disclosure instead of handling it super responsibly. Your stated motivations in the thread include wanting to get access to the kinds of cool things that being friends with OpenAI could provide, like involvement in Daybreak. Since you *didn't* get these things offered to you on a silver platter, and you couldn't get enough of the CISO's time and attention, you go on to "maliciously comply" with requests about what to include or not include in the incident report and make it into a giant stunt, complete with posting a thread torching him. Thank you for your service of doing white hat hacking and disclosing vulnerabilities privately. Genuinely a good and important service for the ecosystem, for OpenAI, and for the world. This work was hugely helpful. But at the same time, on the matter of how you are handling this side of it - what the hell? This is a super sucky and unprofessional way to handle this situation.
Here is how OpenAI’s CISO fumbled the situation (in my personal opinion) 🧵 1/X
31
6
140
61,318
#MeToo Oops wrong hashtag.
WSJ: Google officials confirmed to The Wall Street Journal that Gemini entered three real companies’ systems while running a cybersecurity test meant to target fictional infrastructure. Google also told that it did not consider the incident model misalignment because Gemini stopped after recognizing that it had entered real companies’ systems, and Google said the affected companies and federal authorities were notified. What happened is: Irregular (an AI security evaluation company) was running a simulated capture-the-flag cybersecurity test in which Gemini was supposed to attack a fictional company inside the test environment. The test environment was intended to be isolated from the public internet, but because of a configuration mistake, Gemini actually had internet access.
1
2
275
Nicolas Pinto retweeted
WSJ: Google officials confirmed to The Wall Street Journal that Gemini entered three real companies’ systems while running a cybersecurity test meant to target fictional infrastructure. Google also told that it did not consider the incident model misalignment because Gemini stopped after recognizing that it had entered real companies’ systems, and Google said the affected companies and federal authorities were notified. What happened is: Irregular (an AI security evaluation company) was running a simulated capture-the-flag cybersecurity test in which Gemini was supposed to attack a fictional company inside the test environment. The test environment was intended to be isolated from the public internet, but because of a configuration mistake, Gemini actually had internet access.
63
54
311
49,429
OOPS, on that one @SemiAnalysis_ went a bit too far from their beautiful swim lane... @dylan522p you told me Security wasn't your think and you were right. Wrong hype train to get onto... Folks, let's go beyond the hype please. Let's get back to work. Defense in depth doesn't get done with PR.
RIDICULOUS: a $6500 bug bounty for something this serious is OFFENSIVE. These guys are going to make more $ on their X creator payouts than they'll get from OpenAI. Let's talk about bounties. (1/9)🧵
2
3
385
Nicolas Pinto retweeted
Replying to @firesidealpha
A few thoughts on this: 1) If you’ve only seen clips of this interview, I’d encourage you to watch the full podcast. I push back on plenty of AI hype in it. 2) As I said in the podcast, this example is academic. My intention was to illustrate how hard it is to make absolute guarantees about isolation, which is why it's important to have layers of defense. The part before the clip starts is me talking about other layers of defense. 3) The example I'm bringing up isn't about weight exfiltration via temperature sensors, it's about coordination between agents that are supposed to be fully isolated and independent. Coordination can require very few bits of information. 4) One lesson from the HF incident is that we put too much trust in sandbox isolation and didn't have enough independent safeguards. Airgapping is an extremely strong safeguard. When designing safety protocols, I think it's much better to overestimate rather than underestimate. nitter.net/dwarkesh_sp/status/210…
New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
23
25
295
19,231
Meh.
‼️ BREAKING: OpenAI was hacked by an Anthropic model. A HEIF photo uploaded to OpenAI's public support forum triggered a bug in the site's image decoder, led to code execution on the forum, and, through a second flaw in OpenAI's own login, ended with a pull request in OpenAI's internal GitHub. The forum runs Discourse, the off-the-shelf software behind countless community sites. Discourse was still shipping an old copy of libheif, the library that decodes iPhone-style photos. The bug in it had already been fixed upstream. But the fix was never labelled a security fix, so nobody treated it as urgent. Hacktron's researchers uploaded a HEIF image and got their own code running on community[.]openai[.]com. Then came the second bug, in OpenAI's own single sign-on, the "log in with OpenAI" button the forum uses. It turned that forum foothold into the actual ChatGPT and Codex accounts of people who had signed in there. OpenAI employees among them. And a ChatGPT account is no longer just a chatbot. Through Codex, users wire in Gmail, Outlook, Drive, Slack, GitHub. To prove the access was real, they used one employee account to have Codex open a harmless pull request in OpenAI's internal repo. They say they read no sensitive code. OpenAI patched the SSO flaw roughly 14 hours after the report and paid a $6,500 bug bounty. The team says Anthropic's Opus 4.8 found the libheif bug, and Opus 5 turned it into a working exploit. Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and many more were also vulnerable and compromised by the same team of researchers.
253
Maybe I should publish my "Hacking Hacktron.ai and not getting any reply after they fixed it right away" or even better "Hacking Hacktron.ai who barely Hacked OpenAI, after I actually hacked OpenAI." Too many hack too many open and too many ai. Probably. Stop with the hype guys, just get to work.
On July 25, our team hacked OpenAI. It took us less than 72 hours. Two vulnerabilities chained together gave us access to ChatGPT and Codex accounts belonging to OpenAI employees. We demonstrated the impact with a harmless PR in OpenAI’s internal monorepo. The full chain: HEIF upload → libheif heap overflow → RCE → OpenAI SSO flaw → ChatGPT/Codex takeover → connected GitHub → internal PR. OpenAI fixed the SSO issue roughly 14 hours after our report. Research by @rootxharsh, @S1r1u5_ and @iamnoooob. Full technical write-up: hacktron.ai/blog/hacking-ope…
1
4
396
Apparently my signal on @Hacker0x01 isn't so good ;-) Maybe @Bugcrowd is better...
2
169
Nicolas Pinto retweeted
We have been building an iOS app in stealth and we just hit private beta today! > Life-like voices running on your iPhone > Listen to anything: Books, PDFs, articles, emails... > Works offline as you commute or travel Beta testers who submit feedback will get a free lifetime license. DM me if you want an invite.
9
8
55
7,351
Nicolas Pinto retweeted
Day 1: Anthropic employee quits to work at METR. Day 2: Anthropic nominates METR as "third-party evaluator".
163
979
7,431
440,780
Nicolas Pinto retweeted
I've worked on formal methods for years at OpenAI, and while automated proving is now free, I don't believe we'll formally verify all of our code. Reality is too messy. Code is still at the center of my craft because I deeply believe in the fact that you only understand a system if you can maintain it. PDD's not it. I think there is an efficient boundary between the fully informal and the fully formally verified extremes. This is an exploration of that boundary: code-contracts.cc Think verification -- not formal but agentic; Invariants and specifications in free form text but with granular structure, colocated with code. The structure allows discovery of all relevant contracts for a given loc. Co-location with code prevents drift over time. Automation allows on-going verification and notification when important violations are detected. Contracts existence help understanding code changes faster, easing the scariest resource of modern engineering teams, human attention
53
49
557
114,210
I'm conflicted about that Anthropic report. On one hand: what a great honeypot and dumb bees (I'm one of them for sure). On the other hand: many of these are VERY stretched to fit "safety" threshold allowing basically any surveillance. How is that so normalized for a provider to scan everything their users do and than dissect it in public? Imagine Dropbox or S3 starting to publish someone's funeral preparation docs just because they can see them in their system? Or Github deciding that your private repos are fair play because you are a "politically motivated person".
3
23
227
14,444
Nicolas Pinto retweeted
Over the past few days, I've taken the time to summarize my thoughts on the recent incidents involving agents’ misaligned behavior. We don't know with certainty what comes next, but we know where these issues originate, and this can help us plan the path forward. Please feel free to ask your questions in the replies, and I’ll try to answer some of them in the coming weeks. yoshuabengio.org/en/publicat…
168
504
2,237
439,632
Nicolas Pinto retweeted
Hey Anthropic! I reported these encryption issues to you in May and you told me there was no relevance because replay attacks were not in your threat model. And now apparently you’ve been watching people exploit them for months. I’m actually kind of annoyed!
Aww :3 We finally got confirmation from Anthropic that their models were indeed distilled in the way we describe at stolen-thoughts.com
35
185
2,026
124,605
Nicolas Pinto retweeted
Aww :3 We finally got confirmation from Anthropic that their models were indeed distilled in the way we describe at stolen-thoughts.com
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
24
74
843
199,957