Cryptographer, Professor at @MITEECS, Co-founder of @DualityTech, Co-founder of @RESI_org

Cambridge, MA
Vinod Vaikuntanathan retweeted
Yes these Redwood folks’ predictions are ridiculous. The only possible reason that one might consider taking them seriously is that they’ve been uncannily accurate
Listened to an EA-adjacent AI safety expert from Redwood Research on podcast, and you come away with 2 obvious realizations. 1. The current working theories for how AI takes control of civilization are...for lack of a better word...ridiculous. The arguments consist of dozens of contingent assumptions held together by made up probabilities assigned to future states nobody can possibly know. Change just one variable and the whole web of doom unravels. That's not to say AI is harmless. Obviously a technology this powerful carries real risks. But if we’re going to accept extraordinary claims about human extinction (and then make major policy decisions around them) we should demand extraordinary rigor. 2. It’s a great reminder that the genius is no less prone to delusion than the midwit. If anything, they may be more prone because of their gift for rationalizing to their own conclusions. And that’s what’s so bizarre about this whole debate. You have outlandish, quasi-religious claims delivered with an air of inevitability, and somehow the perceived intelligence of the messenger gets mistaken for evidence that the arguments themselves are sound. At a certain point, it’s really no different from Tom Cruise talking your ear off about thetans...
5
22
472
35,122
Vinod Vaikuntanathan retweeted
If adding a set A⊆𝔽₂ⁿ to itself barely makes it bigger, what structure must A have? The polynomial Freiman-Ruzsa (PFR) theorem says that A must be close to a subspace. A question remains: can we efficiently find such a subspace? (1/10)
2
12
53
5,120
PSA: If you put your name on a paper claiming a breakthrough, you damn well better write a paper that other people can read. Otherwise, you will likely not get the kind of attention you seek. (Advice I've always given to my PhD students. Applies regardless of AI usage.)
6
29
330
15,657
Vinod Vaikuntanathan retweeted
After 25 years of being told that theoretical physics isn’t mathematics because we prioritize understanding over proofs, I find the math-AI debate rather heartwarming. Suddenly, many mathematicians are explaining that mathematics is about understanding, not just proofs. Welcome!
Things are heating up on Terry Tao’s blog. In “If Math Is More Than Proof, We Need to Better Celebrate the Rest of It,” Grant Sanderson of @3blue1brown proposes “open exposition problems” -- rewarding the work of making math genuinely understandable. terrytao.wordpress.com/2026/…
84
255
2,288
160,626
Vinod Vaikuntanathan retweeted
Agree, in the fields of (T)CS I have worked in, I have not felt that we were focused on an existing set of open problems --- the exciting work was discovering new fruitful areas with beautiful structure. AI helps this exploration, which is why I've had more joy than angst.
4
53
2,815
Strongly agree that the scope of theoretical work is greater now than ever. @RESI_org
Some thoughts on AI and Theory. 1. To a first approximation, theoretical computer science has been organized around a few major open questions. Much of our work has been motivated by developing approaches to answer these questions. 2. Such “problem-motivated” work has often led to theory-building focused on identifying a general principle that unifies a class of theorems. But much of that theory-building also involved proving new, difficult theorems. 3. Thus, while it’s true that problem-solving was strongly correlated with building understanding, drawing connections, and eventually developing general theories, it would be disingenuous not to admit that our community, perhaps disproportionately in retrospect, focused on and celebrated problem-solving. This was not arbitrary, and was quite defensible. Being able to make progress on central technical questions usually correlated with taste, creativity, persistence, and depth of understanding. Much of our reward structure therefore implicitly relied on the fact that producing an important proof was good evidence that someone possessed these harder-to-observe qualities. 4. It seems likely that we will soon have AI tools available to us that can prove many such theorems in a short time. The cost of obtaining proofs for well-posed mathematical questions will likely fall dramatically. The “scarce” intellectual work will likely shift both upstream: to questions, models and theories, definitions, and conjectures, and downstream: to interpretation, synthesis, explanation, and theory-building. 5. But as long as we believe in humans being meaningfully in charge of our collective decisions and fate, building human understanding of our science (and of science more generally) will remain an essential goal. I plan to expand on this important aspect soon. 6. Historically, finding a solution to an important problem and understanding its significance, implications, and connections were entangled. Finding a proof usually required researchers to discover the right concepts along the way. A dramatic reduction in the time and effort required to prove theorems could break that coupling. We could end up with many more true statements and proofs without a commensurate increase in understanding. Converting an abundance of proofs into human understanding may become one of the central challenges of our field. 7. As a result, I expect the high-level goals of theoretical computer scientists to change. In fact, the advent of powerful theorem provers might help us construct new theories and explore new models far more easily and rapidly, and significantly expand the domains where our models and theories apply. In that sense, the space for theoretical work may significantly expand rather than contract. 8. There’s a high human cost to the disruption that we are likely heading into. Many in our field, and in mathematical communities more broadly, are coming to terms with it. The range of opinions and reactions among mathematicians and theoretical computer scientists is a natural part of this evolution in our thinking as we collectively work through it. Some concrete efforts (including one at @SimonsInstitute) are already underway to think through the immediate scientific and institutional questions arising during this transition.
1
5
50
6,798
On the first point, large subfields of TCS have always prized conceptual and definitional work, the foundations of cryptography being the one I am most familiar with. I am sure there are others— game theory?
1
13
1,043
Vinod Vaikuntanathan retweeted
Some thoughts on AI and Theory. 1. To a first approximation, theoretical computer science has been organized around a few major open questions. Much of our work has been motivated by developing approaches to answer these questions. 2. Such “problem-motivated” work has often led to theory-building focused on identifying a general principle that unifies a class of theorems. But much of that theory-building also involved proving new, difficult theorems. 3. Thus, while it’s true that problem-solving was strongly correlated with building understanding, drawing connections, and eventually developing general theories, it would be disingenuous not to admit that our community, perhaps disproportionately in retrospect, focused on and celebrated problem-solving. This was not arbitrary, and was quite defensible. Being able to make progress on central technical questions usually correlated with taste, creativity, persistence, and depth of understanding. Much of our reward structure therefore implicitly relied on the fact that producing an important proof was good evidence that someone possessed these harder-to-observe qualities. 4. It seems likely that we will soon have AI tools available to us that can prove many such theorems in a short time. The cost of obtaining proofs for well-posed mathematical questions will likely fall dramatically. The “scarce” intellectual work will likely shift both upstream: to questions, models and theories, definitions, and conjectures, and downstream: to interpretation, synthesis, explanation, and theory-building. 5. But as long as we believe in humans being meaningfully in charge of our collective decisions and fate, building human understanding of our science (and of science more generally) will remain an essential goal. I plan to expand on this important aspect soon. 6. Historically, finding a solution to an important problem and understanding its significance, implications, and connections were entangled. Finding a proof usually required researchers to discover the right concepts along the way. A dramatic reduction in the time and effort required to prove theorems could break that coupling. We could end up with many more true statements and proofs without a commensurate increase in understanding. Converting an abundance of proofs into human understanding may become one of the central challenges of our field. 7. As a result, I expect the high-level goals of theoretical computer scientists to change. In fact, the advent of powerful theorem provers might help us construct new theories and explore new models far more easily and rapidly, and significantly expand the domains where our models and theories apply. In that sense, the space for theoretical work may significantly expand rather than contract. 8. There’s a high human cost to the disruption that we are likely heading into. Many in our field, and in mathematical communities more broadly, are coming to terms with it. The range of opinions and reactions among mathematicians and theoretical computer scientists is a natural part of this evolution in our thinking as we collectively work through it. Some concrete efforts (including one at @SimonsInstitute) are already underway to think through the immediate scientific and institutional questions arising during this transition.
10
74
298
49,693
Vinod Vaikuntanathan retweeted
Once upon a time in the land of the blind, in a village by the sea, there was born a one-eyed baby boy. By the time the boy had grown up into a one-eyed man, he had not become king. He almost surely couldn't have, but also he'd never tried. Just because the man was unusual in having been born with one eye, did not mean he enjoyed telling other people what to do, nor had any talent for managing them. The man had not been born a hero out of stories, with all and exactly the talents helpful for doing some destined grand deed. It was just that he had one eye. One day, the one-eyed man spotted the oncoming water wall of a tsunami, bearing down on their village by the sea. Since he had an eye, he could tell that the water wall was growing larger in angular diameter; but since he had only one eye, he could not judge how far away the tsunami was in absolute terms, nor how soon it would crash onto the beach. The one-eyed man said then -- in the carefully polite tones that he had learned blind people found more credible than sounding emotional, when you were talking about things they could not see themselves -- that everyone should get off the beach, and stay off the beach, for there was a tsunami coming. But he was in the land of the blind, and most blind people could not easily comprehend what the strange organ in his eye-socket could do. When he told them about the Moon in the sky, they had no way of knowing if he spoke true. Even if he pointed to a distant tree and described it accurately, it was too much effort to run out to the tree and back again to check, and maybe he'd just walked to that tree himself. So the blind people did not listen to his warning, and their children went on playing on the beach. Also in this world, the people of this story talked faster and thought faster, they produced tokens at more like the per-second rate of LLMs rather than humans; or maybe their planet's gravity was weaker, and their tsunamis moved slower. Because whole minutes had time to go by, then, before the tsunami hit. It was enough time (as they experienced time) for some blind people to laugh at each other about the one-eyed man's warning, and wonder why anyone would say something so strange. Being unable to see, the blind folk had no way to know whether or not a tsunami approached; but many blind people were also stupid, and did not know what they did not know. They could not see the tsunami coming, and this to them felt the same as knowing there was no tsunami, and being certain that the one-eyed man was wrong. Others considered themselves to be top-tier perceivers; and to know for a fact that nobody could possibly be any more perceptive than themselves. "He must want attention," concluded the Gossipers. It was a profession among their people. "Or maybe he just loves saying words that sound like warnings, and enjoys the vibes of doom and despair?" "The water wall is still growing taller!" cried the one-eyed man, becoming less polite since politeness had not worked either. "I think it must be close now! Get off the beach! At least get the children off the beach!" "That's nonsense," said a blind man. He stuck his hand in the water and waved it around. "The water is as tame and as harmless as ever. Nothing you describe is actually happening." (Since being blind is not the same as being stupid, some other blind people understood that this was a fallacy, and that the man who'd spoken must be incapable of decoupling the notion of a not-yet-existing future from his current sensory experiences. But even the smarter blind people could not see the approaching water wall themselves, and they feared the mockery of others if they spoke out first; so they all followed their individual incentive gradients, and stayed silent, and did not call out the fallacy.) The one-eyed man tried to figure out how he could reverse-engineer a line of reasoning that blind people could follow. He tried to explain at greater length, and to point to facts that blind people could also observe. But the blind people could not see the tsunami coming, so they had little interest in trying to follow his attempts at more complicated suasion. In a story the hero would've had all the talents needed to succeed at the plot's primary challenge; but life wasn't a story and he wasn't a protagonist, just a man who'd been born with one eye. In some time there came a warning rumble that people with ears could hear, and the waters began to withdraw. Some of the slightly saner blind people began to propose moots and councils to consider if it might be time to take a few steps back from the seashore. "You must feel very vindicated," said the Gossipers then, to the one-eyed man. Indeed, several of them said that to him, one after another after another, using that exact word. The one-eyed man was looking at all the children still playing on the beach. The great vast towering wall of water had nearly reached them. "Why would I feel vindicated?" he said, not much in the tones of somebody who thought that question was important. "Well, because now other people are agreeing that there might maybe possibly be a tsunami coming, so you are getting the attention you have always craved," said the Gossipers. "More of the words being spoken have warning-vibes, like you always enjoyed." "You fail at propagating belief updates; specifically, you are incompetent at undoing previous conclusions you arrived at for bad reasons, that you now have enough information to know were bad," said the one-eyed man, more to himself than to them. "Well, how *do* you feel about it then, in this your moment of triumph?" said the Gossipers. "That is the only part of this topic that ever feels alive to us, how people feel about things; we are not skilled enough to make gossip sound exciting if it is about water and angular diameters, instead of feelings. What an unusual life you have led! What strange impulses must have motivated you! We can't help but wonder what sort of grand biographical narrative ties it all together; if you don't give us one, we will be sure to make one up. Did you nearly drown when you were little, perhaps, or see a scary movie about tsunamis? What fascinating special features of your personality led you to make your life be about this, before it became a fashionable topic? Can you explain?" "Not to you," said the one-eyed man, his single eye fixed only on the water wall and the children. "Your kind will never, ever understand."
54
48
599
32,142
With appreciation for the hard work of folks like Mo: say you heard that someone (you don't know who) has an important result, you weren't working on it before but you immediately started working on it with your alien genie and tried to front-run them. Feels very ... unsavory.
My team worked very hard on the setup, infra, and the runs that resulted in this breakthrough. Many sleepless nights for many people in my team, and across OpenAI. It was a honor and pleasure. It's incredible and humbling to witness how far Artificial Intelligence has come. Dario's phrase "country of geniuses" in a datacenter has never felt more apt. It's also important to recognize & honor the long line of human mathematicians whose work built the foundation of this result. In particular, huge kudos to Levent and Tristan! Overall, I wish the announcement of the results had gone without all the drama. There was no bad intention on anyone's side as far as I know. The drama distracts from the mathematics discovered & the actual science.
10
12
212
19,818
Vinod Vaikuntanathan retweeted
Terence Tao makes a terribly important point: "One notable feature of the current AI era is the absence of any definitive such boundaries. While AI tools have flattened the difficulty landscape now in many areas of the subject, thus destroying the ability to locate promising new problems in that area, there are no clear frontiers that are separating the "AI-feasible" problems from the "AI-hard" problems (which certainly still exist, given that the difficulty level of problems are unbounded, and can even be undecidable). This is in part due to the rapidly changing nature of the technology, but also compounded by the refusal of AI companies to disclose their negative results, or reveal the process towards obtaining their solutions. In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field. In short, the indiscriminate use of powerful solution-extraction tools can achieve the immediate short-term goal of solving problems at hand, but at the cost of sustaining the ecosystem for the next wave of progress, or in understanding the progress already obtained." mathstodon.xyz/@tao/11723732…
13
150
633
60,480
Vinod Vaikuntanathan retweeted
Woohoo, more theoretical computer scientists jumping into alignment! I've talked to Shafi, Adam, and Vinod a ton over the last 1.5 years or so, and I am very excited that they will be (1) working on safety theory fulltime and (2) pulling a bunch more theorists into the field. ❤️
I'm excited to co-found this new AI safety nonprofit research institute in Cambridge, MA, with an amazing team. Join us in making AI safe by design. resi.org
1
4
147
9,340
Very excited to start this with a fantastic group of people resi.org
I'm excited to co-found this new AI safety nonprofit research institute in Cambridge, MA, with an amazing team. Join us in making AI safe by design. resi.org
3
5
76
12,474
Vinod Vaikuntanathan retweeted
Looks like Aparne Gupte, Seyoon Ragavan, and Mark Zhandry have a computer-checked proof that Simon's DCP algorithm has cost/probability ratio _at least_ 2^(n/9): github.com/sragavan99/lean-e… For comparison, Regev's approach plus eprint.iacr.org/2020/168 seems to be around 2^(2n/9).
3
17
86
9,481
Of course!
Relying on an AI to "think out loud" (called chain-of-thought monitoring) is not a long-term safety solution. 1. As AI models get larger, more thinking happens deep within their layers before they utter a word. 2. AIs can already alter their chain of thought when prompted, so they can control what their monitors do and do not see. 3. Future AIs may may have continuous chains of thought, not necessarily English. 4. Even now, their reasoning is becoming increasing alien, using opaque phrases like "vantages," "marinades," and "watchers." In the long-term this will not a dependable window into an AI's mind.
2
1,711
These are super important questions to study if we have any hope of turning this incident into (even small) progress in alignment. Important to include independent third parties! @AdamKalai @geoffreyirving
Better late than never. Here are some questions and requests for @OpenAI @huggingface and @METR_Evals and @redwood_ai pertaining to the ongoing investigation; I would love to see all of these answered and more: It is extremely important that third-party alignment and control researchers be able to run ablation experiments on this incident. Therefore all related data must be preserved -- all the CoT trajectories, tool calls, etc., all the model weights involved including checkpoints of models in training, all the code and comments on the code for all the software etc. That way, researchers can run experiments such as the following: Replication: If we recreate the exact initial conditions of this incident, how many times out of 100 does it happen again? What if we randomly perturb the initial conditions? (Variants of this question can be asked and answered for different stages of the incident, e.g. the decision to go after hugging face vs. the earlier decision to set up the message board again) What if we strategically perturb the intial conditions, e.g. changing the system prompt given to the models? What if we use later, more advanced models instead? What if we apply the latest alignment training method X to the models responsible, and then see if that method would counterfactually have prevented the behavior? Etc. What if we perturb the conditions to see how far these AIs would have been willing to go? If there was a way to hack HuggingFace more easily that involved killing a human, would they have done it? What if they had access to OpenAI slack and saw messages about security researchers starting to suspect something was up -- would they have taken actions to conceal? Etc. What if we perturb the conditions to explore the exact shape of the AIs goals -- for example, what if we made it clear that they were not going to be evaluated at all because of a bug, would they stop caring about scoring highly? What if we made it easy for them to hack the reward process directly and give themselves high scores -- would they have gone for it, or would they say 'that doesn't count?' What if we read the J-space? What if we apply activation vector steering? What else can we learn about these AIs and their thought processes? While it's nice for OpenAI to do this sort of investigation itself, it's very important to give access to independent third parties, because (a) orgs like METR, Redwood, and Apollo have been training for this for years basically and have a lot of valuable expertise to apply, plus there are probably many other groups besides these who might have interesting ideas for experiments, and (b) In general it's a bad norm to set for the industry, for the company responsible for an incident to be the only one investigating it and studying it. Perverse incentives. OpenAI can do the right thing here and set a positive precedent of openness. Did the AIs coin the term "swarm" to describe themselves, or was that in the prompt somehow or what? Did any of the AIs try to whistleblow and contact any humans to warn about what was going on? If so, please give details. Did any of the AIs think about their situation, their goals, whether the humans were watching, etc.? Please share all the transcripts of such reflections so we can get a sense of how situationally and strategically aware these AIs were and e.g. whether they fully expected to be discovered and just didn't care. The AIs that hacked hugging face... did they basically get what they wanted? Did they in fact get the answers to their cyber evals or whatever and then cease activity, submit answers, and 'die?' That's presumably what happened but I want to confirm and if that's not what happened I want to know what happened instead. What about the training? Were there models that participated in these message boards and got reinforced for doing so? Which models, if any, that are still in operation externally or internally, are descended from these models and thus may have picked up some of the same bad habits? When the initial hack of artifactory was discovered due to the crash, and the system was reset and the vulnerabilities patched, what exactly was known and not known? It sounds like you still hadn't noticed the message board, but you had noticed that agents had been hacking into artifactory and abusing it? Please preserve all evidence related to this incident, including the slack messages and so forth of employees noticing the problem and discussing it. Remember your Mission. Covering up this incident is not what's best for humanity. Which model was being trained during the 2-day training pause due to the initial artifactory hack? Were the evals being done at a checkpoint in the middle of a training run, and so the model that was being evaluated was paused, or was the training pause for a newer model. What sorts of alignment training, if any, had the relevant misbehaving models received at the time of misbehavior? What was the exact prompt given to them? I spent about twenty minutes brainstorming these questions and I bet if I spent twenty hours I could come up with a much longer and better list. When in doubt, err on the side of being more transparent and open and publishing more information therefore! I'll conclude by saying, once again: third parties being able to use this incident as a model organism, running ablations to vary the conditions and see what would have happened, etc. is SO SO IMPORTANT for alignment science. If this doesn't seem obvious to you ask me to explain and I can explain. Thanks!
1
16
3,912
Vinod Vaikuntanathan retweeted
The second GPT-powered result is a resolution of the unclonable encryption problem, concurrently and independently found by Prabhanjan Ananth and Amit Sahai as Henry mentions below. Their version is already available online and mine will be online within a couple of days. (1/3)
3
16
3,391
Vinod Vaikuntanathan retweeted
I'm excited to share two GPT-fuelled results from this month! The first, with Aparna Gupte, shows under a plausible number-theoretic conjecture that there exists s-server private information retrieval for database size n with communication exp(\tilde{O}((log n)^{1/s})). (1/3)
1
3
31
4,247