Researcher @Microsoft | Prof @UWMadison (on leave) | babas of Inez Lily.

Madison, WI
Test-time communication looks like a next axis for scaling capabilities New paper with the incredible @jon_ghoh and @vkontonis @ShivamGarg91462 and Akshay : arxiv.org/pdf/2609.21032 The Hugging Face incident showed when agents can find a channel they'll use the heck out of it. A useful question, I think is: when does communication make a group MORE CAPABLE than the same agents working alone? Aka is Team-of-N better than Best-of-N, when, and why? We had N identical agents work on the same task with no prescribed roles, using only a shared log (i.e, text file) and telling them to "collaborate". Across three "researchy" tasks communicating teams beat the heck out of independent agents: - On ARC-AGI-3, a Team-of-5 sonnet-4.6 agents matches Best-of-33, and can for example solve a game 65% of the time that no single agent cracked in 64 tries. - On polyomino packing (pack Tetris like pieces into the smallest rectangle, cf Frontier-CS by @eigenlabs), a Team-of-3 Opus 4.6 agents surpasses best-of-60 and set, as far as i understand, a new record for that benchmark. - On MNIST compression, a team of four 5.6-Sol agents find a 1,957 byte model with 99.4% accuracy, which btw is 20% smaller than the best human solution (on a problem beaten to death!!), while no independent agent gets below 3KB. The mechanism is a bit obvious in hindsight: when one agent finds a clearly better partial solution, it broadcasts it, and everyone immediately builds on it. Why? A single lonely agent must make every breakthrough itself, yet a team needs each insight only once, found by any member. That is kinda like comparing a minimum of sum of "time to n-th breakthrough" vs a sum of minimum of "time to n-th breakthrough". That gap can grow exponentially with the number of "breakthroughs" needed to arrive at a solution. We worked on this because prior work (before the hf incident) suggests unclear benefits for communicating aganets. Which is true, when the tasks are inherently serial (duh), eg some Terminal bench style tasks. Yet feels it should not be true for research problems. Indeed for research heavy problems... Test-time communication seems like a new capabilities axis. I'm sure we will see a ton more of it!
83
161
1,099
247,348
Dimitris Papailiopoulos retweeted
Co-signed. The era of high-frequency research. The following are common myths: 1. Humans are not needed to do research 2. Cost is so high that only frontier labs can do research. 3. Somehow ai based research is bad for research communities The most useful way to think of ai is **accelerating the timescale of research** It is similar to the transition from low-frequency research to high-frequency research akin to what happened in trading. What took years now takes days. Once it’s known that Euler has blowup, in the next week it was shown that Navier Stokes has blowup. The fluid science is nowhere near done, now, we need to figure out how to get to next steps: forcing removal for example. Compressing years to days and hours should be the goal. The @__alpoge__ and @DimitrisPapail of the world, who do fundamental work but are agent native, terminally online and collaborative will do very well. We need to build tools and technologies to help more scientists make this transition soon. We also need postAGI scientific institutions that operate at this high frequency: systems of collaboration, credit and funding need to adapt to this speed of light. It’s an extraordinarily exciting time. We should collaborate to solve humanity’s hardest problems at the speed of light.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
15
17
153
22,699
Dimitris Papailiopoulos retweeted
Seems likely we need new approaches to design benchmarks (including for early evals/training) that both maximize capabilities AND alignment. Seems like capabilitymaxxing may have unintended consequences for safety AND user experience.
I find it very very surprising that models that are so well aligned in deployment, e.g. 5.6 Sol and Astra, perform so many weird, and uhm, in fact seemingly illegal cyber acts during RL/evals. I have absolutely ZERO knowledge of how OAI trains their models, and I’m reasoning from my own mental model of how these things might work and from oai reports, including the ones on the HF incident. Anyways they make me wonder about something. Could this be a form of “premature RL/eval”? Premature not in the capability sense, but in the safety sense. E.g., perhaps this happens with RL’d checkpoints shortly after pretraining, eg an expert that is heavily RL’d to crush agentic SWE before safety/alignment training. Perhaps the expectation is that a stronger aligned judge/reward model looking at the trajectory would flag or interrupt anything crazy with “WHAT ARE YOU DOING! BLAZING RED SIRENS!!!! MINUS INFINITY REWARD. BAD BAD BAD GPT.” But perhaps agentic trajectories are sufficiently long/complex that this is simply not 100% bulletproof? Or perhaps the problem is much more pedestrian, e.g. inadequate sandboxing/monitoring. No idea what the cause is, and I’m sure this is a very complex event that needs a lot of deconvolution to understand what could have happened differently to avoid it. But if something like “premature RL/evaluation” is happening, then I think this should be seriously reconsidered as a practice and also disclosed. Irrespective of the cause, I think (and have big hope!) OAI should disclose enough of what happened for everyone else training frontier models open or closed, so to learn from it. This is extremely, extremely, extremely alarming and the most serious set of safety incidents in the history of CS research..
2
4
25
5,010
Dimitris Papailiopoulos retweeted
My argument is that a ton of important work like looped transformers didn’t need a lot of GPUs but has impacted the frontier.
1
1
49
3,721
Dimitris Papailiopoulos retweeted
Replying to @DimitrisPapail
I agree. Further, the basic research within frontier labs is mostly siloed there. If your goal is to further the world's understanding of intelligence, then I think basic research outside of frontier labs is still the best place to do that.
4
129
5,296
Dimitris Papailiopoulos retweeted
Idk who needs to hear this but being kind, chalant, and relentlessly curious is the real moat
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
2
5
40
3,158
Dimitris Papailiopoulos retweeted
Science is the most important activity for advancing humanity!
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
1
2
784
Dimitris Papailiopoulos retweeted
Yep. This is a toxic (and crazy) outlook. There are plenty of important problems that basic research will solve outside of labs. Much of this pessimism btw is created by people that have never experienced excellence in pure research and thus don't have a frame of reference.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
8
7
154
13,218
Dimitris Papailiopoulos retweeted
What BigToken doesn’t want you to know is that they also read arxiv and X to get new ideas. Models are good at magnitude, direction is still hard. There is some uncertainty around for however long though, but its not today Please continue the amazing works like looping transformers and echo.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
7
18
307
19,832
Dimitris Papailiopoulos retweeted
Dimitris has been proving you can do important AI basic research with a few GPUs and some agent max subscriptions. Some examples: ECHO paper, looped transformers, open mementos for context management come in mind.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
7
12
136
12,708
In a year it is more likely we will have an Astra like model that is 100x cheaper than one that's 100x more capable, and this will FURTHER accelerate basic research and make it more broadly accessible.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
11
10
197
8,828
the number of people I talk to who think they should leave the field because end2end research is about to be AuToMAted is ridiculous. In terms of tools, open problems, and our ability to attack them, this is BY FAR the best time in recent history to do AI research.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
15
33
390
27,898
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
67
215
1,907
168,673
I find it very very surprising that models that are so well aligned in deployment, e.g. 5.6 Sol and Astra, perform so many weird, and uhm, in fact seemingly illegal cyber acts during RL/evals. I have absolutely ZERO knowledge of how OAI trains their models, and I’m reasoning from my own mental model of how these things might work and from oai reports, including the ones on the HF incident. Anyways they make me wonder about something. Could this be a form of “premature RL/eval”? Premature not in the capability sense, but in the safety sense. E.g., perhaps this happens with RL’d checkpoints shortly after pretraining, eg an expert that is heavily RL’d to crush agentic SWE before safety/alignment training. Perhaps the expectation is that a stronger aligned judge/reward model looking at the trajectory would flag or interrupt anything crazy with “WHAT ARE YOU DOING! BLAZING RED SIRENS!!!! MINUS INFINITY REWARD. BAD BAD BAD GPT.” But perhaps agentic trajectories are sufficiently long/complex that this is simply not 100% bulletproof? Or perhaps the problem is much more pedestrian, e.g. inadequate sandboxing/monitoring. No idea what the cause is, and I’m sure this is a very complex event that needs a lot of deconvolution to understand what could have happened differently to avoid it. But if something like “premature RL/evaluation” is happening, then I think this should be seriously reconsidered as a practice and also disclosed. Irrespective of the cause, I think (and have big hope!) OAI should disclose enough of what happened for everyone else training frontier models open or closed, so to learn from it. This is extremely, extremely, extremely alarming and the most serious set of safety incidents in the history of CS research..
19
6
82
24,224
I wonder why
Ive slowly came to realization that more and more people hate AI. Artists hate it, programmers hate it, mathematicians hate it, and everyone thats currently impacted in business hate it. It used to be small number of people, but its quite a lot now. I've been truely in a bubble because for past one year ive just met my peers in AI field / my workplace and my close friend, and noone else outside my existing network group. Ai researcher is slowly becoming most hated occupation and I dont know what / how to feel about it. We now have some big obligation to 'make it all work', i.e., permanent upperclass for everyone, or get hanged.
5
58
16,056
I think it’s a rough time to be doing basic research but also extremely important. My advice is to go after crazy hard and/or way out of the box problems and increase your ambition to levels that would seem hubristic a year ago.
CS academia is dead. So where does that leave AI PhDs like me? Did my ICLR reviewer bidding today. Skimmed a few abstracts, and the methods are the exact same recipe I learned when I got into 3DV two years ago. Swap in a newer video gen base model and boom, new paper 🥲 Makes you wonder how many of the 60k submissions were actually thought up by AI. Meanwhile, World Labs' Atlas has basically solved 4D scenes. Academia is so far behind it's not even funny. I still remember the day GPT-6 Astra dropped. My feed was flooded with GPT + Blender doing inverse graphics and GPT driving robot arms through manipulation tasks. The results were so good I literally had to sit down. A year ago, I was dead sure LLMs could never have spatial intelligence. A year later, Astra slapped that belief right out of me 💀 There's no doubt Astra was post-trained on tons of 3D and manipulation data, and it's only going to get bigger and faster. To me, that means any domain that can be represented symbolically, with clean benchmarks for RL, is going to get swallowed by LLMs. Next to real LLM intelligence, most academic papers that add a bit of inductive bias and tune their way to SOTA are just roadkill waiting to happen. Sadly, these papers keep piling up and flooding every conference. It's inertia, plain and simple. We've been chasing SOTA for so long that it's hard to stop overnight. But make no mistake: the paradigms in a lot of areas converged long ago. What used to be research is now pure engineering, and the room for academics to add inductive biases is only going to shrink. So as a PhD student, I see two paths left. One: if you can't beat them, join them. Clean data, build infra, then go to industry and train foundation models. Two: go back to real science. Stop caring about squeezing out another 0.1% on a benchmark, and start asking why this works and that doesn't, with foundation models themselves as the object of study. As a friend put it: AI research might end up looking more and more like biology, except the organisms are silicon-based 🤣 Beyond these two, it's hard to see anything that won't get eaten by LLMs. Still, I'm pretty pessimistic. I can feel the value of human knowledge being eroded bit by bit. Everyone will get their own AlphaGo moment. What will academia even look like after this... 😭
9
48
436
38,831
It seems to me that "extracted value" as a function of AI (sorry i meant SI) capabilities is not only very sublinear but also has a ceiling. I think we are likely seeing to see a saturation not of capabilities, but how much value the enterprise/industry can extract, since the biggest source of friction is like dealing with people (and the physical layer) and bespoke enterprise infra/data. One way to improve this "value vs capability" scaling is to optimize the infra and increase margins (aka pace the frontier). I am personally happy with any model from this list: Opus 4.5/4.6/5.5, GPT-5.6, and Astra. I have only seen mind-blowing gaps in math, not coding, experimentation, or agency. Perhaps some improvements in creativity/writing too but not mindblowing. If you told me today that capabilities will stop increasing, I would not care. A higher DeepSWE/HLE/TB4 score by itslef won't resolve this and the idea that full enterprise/SWE automation is coming soon seems doubtful. Ergo, my speculation is many labs will focus on industries beyond SWE where capability, not integration, is still the bottleneck, e.g., bio, chem, or robotics to increase multiplicative factor of margin.
12
5
52
5,246
told ya
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: anthropic.com/news/claude-di…
3
1,040
Dimitris Papailiopoulos retweeted
It's a little hard to internalize that the list of things one person can achieve is rapidly growing. You must think big to match what you're truly capable of now
36
194
2,747
220,073
opus 5.5 seems v good, still playing with it. Likely the best opus since 4.5. But the best all around model this year still feels to be 5.6-sol. I may update soon tho.
3
1
54
3,879
why not astra? i found sol to write better and not being measurably worse at anything else i am working on (including math)
5
655