CS prof at Penn. Amazon Scholar at AWS. Author of The Ethical Algorithm (w/ Michael Kearns). I study machine learning, privacy, game theory, and uncertainty.

Philadelphia, PA
Pinned Tweet
A new semester, and the first lecture is in the books for my class on the "Mathematical Foundations of AI Alignment". aaroth.github.io/cis-7000-ai… What does that mean? Good question. We have about a semester in which to figure it out.
11
100
623
51,451
An annoying part of the twitter/bluesky discussion of whether LLMs are "only" "next token predictor"/stochastic parrots is that "next token predictor" is contentless - all mappings from inputs to output strings can be factored into a sequence of "next token" distributions.
6
7
93
3,659
In this sense you, or any other system that can output text can be written as a "next token predictor". This isn't making any statement about whether brains function similarly to LLMs or not at an internal level, its just a boring syntactic fact about probability distributions.
3
2
19
1,065
The interesting question a decade ago about representing output distributions as next token distributions was computational: maybe this representation would introduce computational difficulties that would make it hard to learn useful behaviors. It turned out it didn't!
2
23
829
AI tools are very useful for mathematical research. But at present, producing good work requires a labor intensive step that I have been calling "deslopping" - making the thing readable. Please do not skip this step; even if the theorem is great an unreadable paper is worthless.
12
27
213
24,600
Aaron Roth retweeted
looking forward to this launch event, AI Alignment & Safety, Oct 19th @ Princeton!!
11
47
3,147
Aaron Roth retweeted
my suggestion is more extreme than most I'm seeing here in the comments... but desperate times call for desperate measures, and as the 60k ICLR submissions show, the ML community has been too slow to rise to the challenge. the gist of my proposal is to institute two parallel submission tracks; let's call them Track A and Track B. at submission time, authors choose which track. in Track B, papers are reviewed and judged entirely by AI. This is not to say that it's an accept-all, AI-slop track: the conference still sets standards for acceptance, and papers that don't meet them are rejected. The difference is that those standards are enforced by an AI review process, not by human reviewers. each conference can design this process however it likes, based on the principles it values and acceptance criteria set by its human steering committee. It could choose to mimic existing reviewing practices (blind reviews, multiple reviewers, discussion, AC/meta-review), or invent something entirely different from scratch. I expect this process to improve empirically over time. in fact, the community can actively research better methods for AI review, test where they fail, and iterate from conference to conference. this process itself should be public and scrutible; the new open review, of you'd like. crucially, in Track B, there is no limit on the number of submissions. in Track A, on the other hand, papers are reviewed and judged by human reviewers, much as they are today. But there is a hard cap on how many papers an individual can submit to this track, to keep the demand for human attention under control. I have in mind a cap much smaller than what people are currently churning ou, say, something on the order of 1–3 papers per author per conference. The exact mechanism is a design question; the important part is that access to human review is deliberately scarce. are you a PI whose students typically produce 10–15 papers per conference? great. think hard about which ones would really benefit from human review, or which contributions you especially want the community to pay attention to. submit those to Track A, the rest to Track B. in the longer run, maybe this actually aligns incentives much better. e.g. two talented students could be incentived to join forces on one really ambitious project with their PI as author, instead of producing six incremental works between them without their PI as author. I don't think it's controversial to say the community would benefit from fewer, more substantial research contributions, and I think this has been true for years. AI agents coming for our peer-review process may simply be the tsunami strong enough to finally get enough stakeholders onboard to change the status quo.
I'm a Program Chair for ICML 2027. We'd love any creative suggestions for a great conference! How to manage the explosion of AI slop and insane submission growth? How to manage reviewing? How to incentivize high quality creative work? How to reduce bureaucracy and overhead?
16
17
153
26,662
Aaron Roth retweeted
Replying to @boazbaraktcs
An interesting interaction between STOC's new arxiv policy and arXiv's new one-strike policy is that if you violate the one-strike policy, you seem to get a lifetime ban from STOC. Not necessary unreasonable but I wonder if intentional.
Replying to @tdietterich
The penalty is a 1-year ban from arXiv followed by the requirement that subsequent arXiv submissions must first be accepted at a reputable peer-reviewed venue. 4/
1
2
22
3,797
Codex and Claude Code have a neat auto approve feature where 1) it doesn't ask for permissions, but 2) you get to feel safe. Well --- do you? The premise is it is asking an agent for permission. But if you don't trust the driver agent, why should you trust the review agent?
10
14
76
7,421
But despite being weaker it is not guaranteed to be satisfied. Still, we find some preliminary evidence of non-trivial safety/completeness tradeoffs among real reward models, and these effects come from coallitional rather than individual alignment.
1
1
3
413
Our paper is here: arxiv.org/abs/2609.15803v1 comments welcome! This is joint work with the wonderful @sikatasengupta @natalie_collina and @SurbhiGoel_
2
3
455
A world without open problems Here are some that fell today: K-server: arxiv.org/abs/2609.15979 Matroid Secretary: arxiv.org/abs/2609.14555v1 Matrix Spencer: arxiv.org/abs/2609.15025 (Well Matrix Spencer was maybe also a few weeks ago, but who's counting? arxiv.org/abs/2608.28816 )
We are clearly moving to a world that will be without open problems, but this doesn't mean math is going away. Interesting work will correspond to discovering new questions and coherent theoretical frameworks. Well defined agreed-to-be-interesting problems won't last long.
5
35
233
41,670
Aaron Roth retweeted
Agree, in the fields of (T)CS I have worked in, I have not felt that we were focused on an existing set of open problems --- the exciting work was discovering new fruitful areas with beautiful structure. AI helps this exploration, which is why I've had more joy than angst.
4
53
2,815
I sometimes see people trying to use complexity theory to argue that building "true" AI is impossible. I find this unreasonably annoying. It requires ignoring what is in front of your face and it ignores that worst-case complexity has been an awful guide in machine learning.
8
14
134
12,678
Applying this kind of argument to more general intelligence would have been wrong a decade ago but now requires ignoring the obvious success of AI and particularly the reasoning models from the last few years. A decade ago we could have said it denies computationalism:
1
5
831
i.e. that the laws of computation bind human cognition just as much as they bind silicon. This is merely a conjecture --- maybe its not true and we have some non-computational secret sauce. But now making this argument requires ignoring evidence in silico (and in browser tab)
1
3
768