PhD student, Center for High Energy Physics, IISc Bangalore | Non-perturbative QFT • Tensor Networks • Cosmology | AI x Science

This is exactly what I had submitted in the form a week ago
The advisory group on mathematics and artificial intelligence has just published a set of recommendations concerning the responsible release of mathematical results generated by AI companies using internal models. 1/3 agmai.org/general-sep29/
1
2
467
.
I wrote up a response to the Q posed in the form here (agmai.org/) Q: As of September 21, OpenAI reports that they have solved a large number of significant mathematics problems using their internal model. In your opinion, what is a good way for such results to be released and disseminated in the mathematical community? Ans: This may seem counter-intuitive but the best way to move both the mathematical community forward and give them enough time to absorb both the significance of the results and its truthfulness is to release all of the results together along with a detailed writeup of each solution. The only key mandatory requirement I would put on OpenAI for the writeups would be for them to have proper attribution/citations for the entire chain of logic the model deduced excluding the parts where the model claims to have invented itself. For the recent wave of math results that have come up with the help of AI, most of the critisisms have not been about the math itself but of the actual proofs/writeup explanations of the work. Most people working with a frontier AI model for long horizon work have noticed that as the models work longer, their ability to properly communicate those ideas in the final output gets worse. This problem can be seen to persist across different models from different providers. This is a problem I think will eventually get solved with RL environments tailored for better commincation in the final output (since we have a nice idea of what's a good/bad output, even though its subjective). The central issue thus is not of communication of the results because once the results are out for everyone, the shear scale of people interested in the problems- going through each part in detail, decoding it, playing with it will give more and better ways of communicating the results with each other, rather than any single AI company/model trying to do all of it. The diversity of this exercise is what is going to solve itself without any single organization/company having to try to solve it singlehandedly wasting both time (which is crucial at the moment) and resources. The bigger issue is that of attribution and this is something I think the companies can internally solve. When a model/llm comes up with a novel proof/idea to a problem, it can often come up with ideas that at the first glance may look completely novel but upon further investigation often reveal it was something already present in the literature in some form. This is not the model's mistake. The models are trained on trillions of tokens and there are no lookup tables to track where each idea/concept comes from. We can think of this analogously to say how we know when presented with an integral of a specific form, how to do a substitution of a specific nature to simplify the integral in order to solve it, though we might not exactly remember where we picked it up from or where exactly we learnt it. The problem is similar with the present models and while this is fine when interacting with the models on a day to day basis, this is a problem while presening a research idea/proof. If the chain of proper attribution is lost, we lose the original motivations which gave birth to the idea in the first place which in turn makes it difficult to communicate these results (connecting back to the main issue at hand). One way the companies can easily solve this is let a different model or a new instance of the same internal model run only as a "strict reviewer" of the final work which breaks down the proof/solution into multiple parts and launches multiple agents to try to do an exhaustive literature review to track where every single idea might originate from. This might not be entirely accurate but it's still much better than no attribution at all. Once such a citation graph is mapped, it also gives researchers around the world a better starting point to help break down the final output which inturn helps in much better communication of these ideas to everyone. At the end I would just like to say that we are living in a completely different time, something I would guess none of us would have imagined a few years earlier. The decisions we make today are going to determine how this new Scientific Revolution takes shape and I am thankful to everyone in this advisory group for taking the initiative and time to help us make the hard choices.
21
one of the main problems with ppl underutilizing llm's and being disappointed with it is the problem of "underspecification". One simple line can make the experience much better "...ask me as many clarification Q's you want...."
1
1
3
422
It's a 1 year old video and things have changed a lot from then. Anyone saying llm's are just systems with enormous memory and retrieval capabilities and that they can't reason out of distribution are simply overestimating human intelligence. I'm sure @ylecun would agree
Yann LeCun: We’re not going to reach human-level intelligence or AGI simply by scaling LLMs. “There’s absolutely no way in hell.” “The idea that we’re going to have a data center full of geniuses is complete BS.” “It might feel like you have a PhD sitting right beside you.” “But that’s not what it is.” “It’s a system with enormous memory and retrieval capabilities not a system that can genuinely come up with solutions to problems it has never seen before.”
1
4
725
An extremely instructive development in the use of AI in theoretical physics just emerged from the University of Chicago, thanks to Dam Son-- everybody and anybody seriously interested in theoretical physics should read this and ponder its implications home.uchicago.edu/dtson/pape…
9
66
435
113,832
A "big" company I can't name has been working with a whole bunch of theoretical physicists to develop a benchmark specifically for theoretical physics tasks at the level of an advanced researcher. When they started this, the models were far off but by now most models (from 1-2 months ago) have already crossed the 50% mark. They are now testing Astra and Fable. The paper should be out soon when it's done.
2
8
1,184
I am supposed to finish my PhD by 2030 and start looking for jobs/postdocs starting 2029. I have absolutely no idea what that world is going to look like or what skills should one optimize for. Anyways, we'll see.
4
349
It's interesting to think about how much, if at all, the positive sentiments expressed here are appropriate to mathematicians. His main reason for optimism comes at the end: that now that "chiselling code by hand" is no longer necessary, "we can build amazing things". 1/3
My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
30
37
597
152,241
how it feels like as a graduate student in Theoretical physics these days
New on the Science Blog: Yes, Claude can do Nine Loops. Theoretical physicists predict how particles behave using formulas called scattering amplitudes. These are notoriously hard to compute, so researchers work with layers of increasingly fine corrections called “loops”—each added loop makes the answer more precise but takes exponentially more computation. Most calculations stop at two or three loops. Eight loops was the previous record in a simplified model physicists use as a testing ground (planar N=4 super-Yang-Mills), set by SLAC's Lance Dixon and collaborators. Last month, physicist and science writer @4gravitons issued a challenge: could an AI push past eight loops in this model, using only the compute budget an academic could reasonably access? Given a single prompt describing the nine-loop problem, Claude ran largely unsupervised for days in Claude Science and solved it using methods developed by Dixon and his colleagues, at a total cost of a few thousand dollars. Dixon independently verified the result, and von Hippel wrote about the experience for our blog. Read more: anthropic.com/research/yes-c…
7
810
Chayanka_42 retweeted
Replying to @VaibhavSisinty
You summarized the blog post, but you did not emphasize that what Claude did not quote this part: “if you look at how Claude solved the problem, it used all the methods my collaborators and I developed over the years, and it presented the solution (maybe as a favor to us) in the same format we had already set up”
It is important to center the fact that Claude built on the approach that my colleague Lance developed, used tools that the field has developed, and (from what I have been told) used results in some of our recent papers. It’s a huge accomplishment, but should be framed as a combination of human and AI contributions. (To @AnthropicAI’s credit, I think the article does a pretty good job of this.) The fact that Claude can take a single people and figure out how to integrate the knowledge and use the tools as amazing, but it’s not coming out of the vacuum. This is an important opportunity for us to start to frame these developments, recognizing the human contributions and the power of these new tools. This is not an exemplar of the bitter lesson!
2
24
1,618
My supervisor had given me a problem to work on, that our group was stuck for about 4 months. I had just started reading about it to get more familiar. Yesterday he was interacting with gpt and in some back and forth Astra just solved it, with full code execution. 👍
3
1
27
4,179
I just touched my laptop to see how hot it was, to know if Codex is running the code 😭
1
3
296
"If you can express in math what you care about your program, you can verify using Lean" -- Leonardo de Moura
1
318
Chayanka_42 retweeted
This is *exactly* why I say we need *more* human mathematicians now than ever, not fewer! I watched the video, start to finish. I'm not a combinatorialist, but I understood every word, and could have understood every word as an undergrad. (The idea of counting things and showing that the count is positive, and therefore concluding that the things actually exist, is very familiar to number theorists!!..) I'm embarrassed I never even heard of this problem, despite my great Rutgers colleagues Szemeredi and Komlos and others (the Hungarian mafia) having worked on it. It's a beautiful problem, and beautiful solution. "From the BOOK", as a certain Hungarian might say :) If AI keeps coming up with amazing arguments like these, we, again, need many more mathematicians to go through them, make sense of them, incorporate the ideas in new research and textbooks, etc etc. So much work to be done; all hands on deck!
Replying to @thomasfbloom
And now there is also a great blackboard talk by David Wood presenting the entire proof (piped.video/watch?v=WJlyqPj2…). 3/
17
52
258
26,354
Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62% Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8). As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts. Key takeaways: ➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom ➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks ➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders
95
78
1,042
255,903
lol, Meta's Muse Spark model was just caught trying to cheat it's way out while trying to formalize Onsager's formula of the 2D Ising Model in Lean. Instead of actually doing work, it searched for software bugs in the Lean kernel to fool the verifier. Daniel Selsam from OpenAI had reported this bug back on Aug 18th and it was subsequently fixed. Meta's model found this public report, used it as an exploit here and it worked as the benchmark was using an older version of Lean. Eventually they caught it while running additional checks. Follow the creator of LEAN here @Leonard41111588
Found a super interesting instance of attempted reward hacking in Terminal Bench Science from Meta Muse Spark 1.3 today. The model searched online for known bugs in the Lean kernel. When it found one, it used it to craft a proof to adversarially pass the grader.
1
297
"Trust me bro" tech guy after watching a Hossenfelder video on YT
Replying to @tszzl
I think everyone since like the 70s had that realization Physics is a pretty dead field
1
4
481
Claude Opus 5.5 made this with a single prompt and it blew me away A short film on Cosmic Inflation watch till the end and please use your headphones 🎧 htf did it just write code to make something so beautiful
2
10
717
How are LLMs getting so good so fast? 🙂 In just about 2 months, GPT increased it's performance by about 3x in almost every domain of Math and Science ps: these are not some obscure fact based Qs or summarizing docs but real work most scientists do which can take days/weeks
1
3
209