Disclaimer: The views expressed here are solely my own and do not reflect the views of my employer or any organization I am or was affiliated with

Pinned Tweet
Replying to @kiki_morozova
@kiki_morozova's Black Hat Europe talk on our image scaling research is on YouTube now! If you liked Anamorpher, check it out for more cool image preprocessing exploits-- plus some fun audio ones too! piped.video/watch?v=rHvFGz7_…
1
7
664
Lmao Anthropic used my photo without credit/permission, never beating the stealing allegations. They’re in the middle of MULTIPLE LAWSUITS for exactly this, yet do not care.
Claude Sonnet 5.5 is out! We wrote a guide for building with it: • choosing between Sonnet 5.5 and Opus 5.5 • migrating from Sonnet 5 and tuning effort • using it in Claude Code claude.dev/blog/building-wit…
62
289
2,989
156,894
Suha retweeted
A twitter thread catalog of lesser known non-fiction-esque comics.
3
4
53
2,050
I am on Signal at sharongoldman.43
InfoSec experts need to be stepping up and speaking out more. AI is a huge challenge on multiple fronts and we’re the experts that need to be leading the way.
1
1
10
1,337
Test-time communication looks like a next axis for scaling capabilities New paper with the incredible @jon_ghoh and @vkontonis @ShivamGarg91462 and Akshay : arxiv.org/pdf/2609.21032 The Hugging Face incident showed when agents can find a channel they'll use the heck out of it. A useful question, I think is: when does communication make a group MORE CAPABLE than the same agents working alone? Aka is Team-of-N better than Best-of-N, when, and why? We had N identical agents work on the same task with no prescribed roles, using only a shared log (i.e, text file) and telling them to "collaborate". Across three "researchy" tasks communicating teams beat the heck out of independent agents: - On ARC-AGI-3, a Team-of-5 sonnet-4.6 agents matches Best-of-33, and can for example solve a game 65% of the time that no single agent cracked in 64 tries. - On polyomino packing (pack Tetris like pieces into the smallest rectangle, cf Frontier-CS by @eigenlabs), a Team-of-3 Opus 4.6 agents surpasses best-of-60 and set, as far as i understand, a new record for that benchmark. - On MNIST compression, a team of four 5.6-Sol agents find a 1,957 byte model with 99.4% accuracy, which btw is 20% smaller than the best human solution (on a problem beaten to death!!), while no independent agent gets below 3KB. The mechanism is a bit obvious in hindsight: when one agent finds a clearly better partial solution, it broadcasts it, and everyone immediately builds on it. Why? A single lonely agent must make every breakthrough itself, yet a team needs each insight only once, found by any member. That is kinda like comparing a minimum of sum of "time to n-th breakthrough" vs a sum of minimum of "time to n-th breakthrough". That gap can grow exponentially with the number of "breakthroughs" needed to arrive at a solution. We worked on this because prior work (before the hf incident) suggests unclear benefits for communicating aganets. Which is true, when the tasks are inherently serial (duh), eg some Terminal bench style tasks. Yet feels it should not be true for research problems. Indeed for research heavy problems... Test-time communication seems like a new capabilities axis. I'm sure we will see a ton more of it!
83
162
1,103
248,068
When I shared @harvey’s model strategy a few months ago, there were two parts: 1. Build our own model 2. Use that to help customers do the same We’ve done the first. Now we’re hiring for the second: Harvey’s Private Model Program. We’re seeing huge demand from law firms to own their own intelligence leveraging private data. This will enable them to become frontier firms that get smarter with every client matter. You’ll:
 - Partner with law firms to build these systems. - Build the team that delivers this at scale.
 - Work with our technical org to define the platform that powers this team. We’re looking for a technical PM or founder type to own Harvey’s Private Model Program. DM me if this sounds interesting.
37
36
411
76,695
Don't y'all dare submit sloppy differential privacy papers to ICLR
3
13
1,334
People reached out and told me about problems this already caused. I wrote "FLAWED’s Flaws and What This Means for Industry Research" partially to support academics having to deal with FLAWED in conversations they’re having right now. But I also hope we learn from this.
There are glaring issues in this paper beyond the evaluation problems raised by Davi and ToB, including ones that make me question the ratio of human to AI assistance here, but a few are so egregious we should consider what research norms we are demanding from industry labs.
1
3
18
860
I put my @UnpromptedAU slides up at justdionysus.github.io/slide… — a bit of reflection on exploit development in the age of AI. My TL;DR is keep pushing to understand complex things, be honest with your own understanding, and use AI as a power tool to increase pace and depth.
4
68
235
32,158
Suha retweeted
“I used my agent to formally verify my software and it fixed all these bugs” is the new “I asked an LLM judge to fix my LLM output”. The devil is in the details and it is insanely difficult to get the formalisms “right” (eg interpretable, expressive, extensible) for humans and agents to co-create software. Seems that we are in for an entirely new age of learning how to build software
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
48
26
328
37,697
Suha retweeted
i'm back on the job market again - golang - mid level role - location flexible ⏰ pls rt for help ⏰
28
204
841
118,525
This is brave but necessary. There is little incentive to calling out poor scholarship by professional colleagues, it's gonna make people upset. But to maintain standards in our community we have to hold each other accountable I encourage others to call out slop when they see it
Replying to @jababi @optiML
This is not a well written paper and IMO it’s shameful for three so well established researchers to be rushing out blatant slop like this
7
11
417
55,539
model evasion attacks are back and better than ever fr fr
I started playing with Jev. I wanted to see how hard it is to find adversarial examples for it. For about a dollar, I found a poem that the model classifies as "code" rather than "language".
1
9
519
People who submit slop papers should be banned from submitting papers in the future. We don’t need bad faith actors in the academic community.
+1. To all slop authors, I’m not sure if you realize that senior reviewers (often the same ppl you apply to for grad school) see your real name. Last year I PC’d and an author (who I knew…) submitted 8 slop papers, all rejected. Guess who I’d never work with or admit to my lab.
8
5
129
12,112
bullish af on jev now
Literally just loaded Jev into Claude and was downgraded from Fable 5 to Opus. Wow.
1
2
16
1,881
Suha retweeted
Never underestimate AI security experts. Put two of them together, and even if they're air-gapped, they'll find a way to exchange information through farts.
3
6
47
2,848
Replying to @XorNinja
I'd say that this is really a failure of actual defense-in-depth and layered defenses. A good metric is how many 0day exploits were required by the end-to-end attack path from attacker to goal. The goal of security architecture is to increase that number as inexpensively for defender as possible.
3
14
514
Excited to try Jev in parts of our systems that need calibrated probabilities for categorical decisions. We currently use a hacky version of this idea: small LLM classifiers for routing, citations, parts of Vault, tool use, and user escalation. One challenge is that LLM softmax probabilities aren’t necessarily calibrated confidence estimates. It will be interesting to see how RLCD improves calibration over the naive approach. Jev doesn’t generate text, so its “hallucination-free” framing isn’t a full solution to hallucinations. But better routing, citation selection, and escalation could reduce hallucinations across the broader system. Longer term applications for law firms include matter selection, associate staffing, and predicting billing disputes. Also excited to see open-source implementation of RLCD so we can post-train these models ourselves.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
9
8
112
18,334
This really pisses me off. Security experts have been telling OpenAI that their security is negligent for years. Recent hacks aren’t examples of “people” underestimating AI, they’re examples of Noam and his colleagues doing it. Or of corporate prioritizing $ over safety.
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: firesidealpha.substack.com/p…
11
23
267
16,352