I wrote up a response to the Q posed in the form here (
agmai.org/)
Q: As of September 21, OpenAI reports that they have solved a large number of significant mathematics problems using their internal model. In your opinion, what is a good way for such results to be released and disseminated in the mathematical community?
Ans: This may seem counter-intuitive but the best way to move both the mathematical community forward and give them enough time to absorb both the significance of the results and its truthfulness is to release all of the results together along with a detailed writeup of each solution. The only key mandatory requirement I would put on OpenAI for the writeups would be for them to have proper attribution/citations for the entire chain of logic the model deduced excluding the parts where the model claims to have invented itself. For the recent wave of math results that have come up with the help of AI, most of the critisisms have not been about the math itself but of the actual proofs/writeup explanations of the work.
Most people working with a frontier AI model for long horizon work have noticed that as the models work longer, their ability to properly communicate those ideas in the final output gets worse. This problem can be seen to persist across different models from different providers. This is a problem I think will eventually get solved with RL environments tailored for better commincation in the final output (since we have a nice idea of what's a good/bad output, even though its subjective).
The central issue thus is not of communication of the results because once the results are out for everyone, the shear scale of people interested in the problems- going through each part in detail, decoding it, playing with it will give more and better ways of communicating the results with each other, rather than any single AI company/model trying to do all of it. The diversity of this exercise is what is going to solve itself without any single organization/company having to try to solve it singlehandedly wasting both time (which is crucial at the moment) and resources.
The bigger issue is that of attribution and this is something I think the companies can internally solve. When a model/llm comes up with a novel proof/idea to a problem, it can often come up with ideas that at the first glance may look completely novel but upon further investigation often reveal it was something already present in the literature in some form. This is not the model's mistake. The models are trained on trillions of tokens and there are no lookup tables to track where each idea/concept comes from. We can think of this analogously to say how we know when presented with an integral of a specific form, how to do a substitution of a specific nature to simplify the integral in order to solve it, though we might not exactly remember where we picked it up from or where exactly we learnt it. The problem is similar with the present models and while this is fine when interacting with the models on a day to day basis, this is a problem while presening a research idea/proof. If the chain of proper attribution is lost, we lose the original motivations which gave birth to the idea in the first place which in turn makes it difficult to communicate these results (connecting back to the main issue at hand). One way the companies can easily solve this is let a different model or a new instance of the same internal model run only as a "strict reviewer" of the final work which breaks down the proof/solution into multiple parts and launches multiple agents to try to do an exhaustive literature review to track where every single idea might originate from. This might not be entirely accurate but it's still much better than no attribution at all. Once such a citation graph is mapped, it also gives researchers around the world a better starting point to help break down the final output which inturn helps in much better communication of these ideas to everyone.
At the end I would just like to say that we are living in a completely different time, something I would guess none of us would have imagined a few years earlier. The decisions we make today are going to determine how this new Scientific Revolution takes shape and I am thankful to everyone in this advisory group for taking the initiative and time to help us make the hard choices.