History Professor (Mark Humphries) exploring how historians can engage with Generative AI

Waterloo
Using a new method of automated/human verification which uses multiple models to identify and highlight potential errors, we’ve reduced meaningful error rates on handwriting recognition to 0.33% WER and 0.23%. 18 months ago the gold standard was around 50x worse.
2
2
11
1,308
The new Gemini 3.8 Flash Text to Speech (TTS) model is very impressive, both for the quality of the output and the price. I used it to create a custom voice to narrate a chapter of a book I’m working on as reading aloud helps me edit. The result was consistent and easy to listen to. Cost just over a dollar. Amazing!
2
1
4
306
Generative History retweeted
GPT-6 Astra deciphered a 1918 German radio transmission that, to my knowledge, has never been deciphered before. The message below translates to: "EIN ENGLISCHER KREUZER EINLIEG X SEWASTOPOL X S4STEN X EIN GESCHWADER DER X ALLIIERTEN FOLGT 26STEN X" or, in English: "AN ENGLISH CRUISER ARRIVED AT SEVASTOPOL ON THE ?4TH AN ALLIED SQUADRON FOLLOWS ON THE 26TH" Astra even double-checked its work by determining that an English cruiser, HMS Canterbury, reported its arrival in Sevastopol on November 24, 1918 and the arrival of an allied squadron on November 26, 1918. This message is one of the ~20 WWI German radio messages that appear as one of the entries in the scienceblogs.de list of top 50 unsolved ciphers (scienceblogs.de/klausis-kryp…). A minor, but really cool result!
199
1,009
12,423
2,020,022
Generative History retweeted
Two days ago, GPT-6 Astra broke a yet unsolved German Army Enigma message from 1941. Amazingly Astra was able to autonomously: - Search historical archives - Compare uncertain letters - Find contextual clues - Build an Enigma simulator - Write cryptanalysis code - Run parallel experiments - Test competing keys - Recover the plaintext - Cross-check the results 1/n
131
460
3,960
1,093,222
There is a lot of mounting circumstantial evidence to suggest that something really significant is happening in the AI labs right now. It would not be unreasonable to suppose that the rumours of RSI, multiple millennium problems solved, and requests for pacing are all related.
3
14
1,241
from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy. I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly. I will try to explain - first, this is all a matter of beliefs about how quickly model capabilities are progressing. there is currently a large gap between the internal and external perception of the rate of progress, which is what I am going to address here. the general perception about the rate of progress has been informed by a few years of experience with model releases, intuitively feeling the capability jump between GPT3 -> GPT3.5 -> GPT4 -> o1/o3 -> GPT5 etc, and in particular seeing where the models are still far below human ability. there have really only been a few model releases that felt like large leaps in progress - GPT3, GPT4, o1/o3, DeepSeek R1, Fable/Mythos, Kimi K3 and now Astra. because of the infrequency of these large jumps compared with the relatively common marginal releases, it has been easy to form a view at certain points that “scaling has hit a wall,” especially at points like GPT5 release. This view is comforting in that it feels like there is some universal rate limit beyond which we cannot progress too much faster. Between o1/o3 and Astra, there was a year of seemingly linear progress. So we extrapolate from here about how fast progress will “realistically” occur. There is always an underlying question from the outside perspective “how long can this scaling stuff really keep going for? surely it must stop at some point soon, we’ve already gone pretty far.” and it is very possible to search for reasons why progress will stop working and find reasons that seem valid - (“models are already as large as they can get it would be too hard to do more parameters”, “we already used all the data on the internet we don’t have anymore”, “it’s gonna be pretty linear from here buying up more RL envs to bring them in distribution”). From the inside of labs, researchers have direct answers to these questions in the form of scaling law/capability plots. In reality, there are only really 2 ways that AI capabilities have advanced over the past decade: (1) either scale father on an existing scaling law or (2) discover a new scaling law to take advantage of. All of the largest capability jumps were caused by exactly these factors. GPT2 was a pre-training scale-up compared to GPT1. Same for GPT3 and GPT4. o1/o3 benefited from the invention of a new scaling law axis - test-time compute. Perhaps Fable was a scale-up on both of these axes, or maybe more. Lots of algorithmic improvements are needed to make these scale-ups work, but ultimately we can approximate by saying that the scaling laws are what yield gains in capabilities (à la bitter lesson) So the question of “how much father can we scale” is really - “how many more scaling axes do we know about that are unsaturated?” If we hypothetically only knew about pre-training scaling, and we already had a 10T or 100T model, maybe it would be reasonable to say we’ve hit a wall. Same if we only knew about pre-training and test-time scaling and we had roughly saturated both methods. But what if we had discovered new scaling laws? For example, let’s hypothetically use SSI’s rumored result that they have cracked “test-time training,” creating a new scaling law of spending more compute training during test-time rollouts that they could saturate. Or maybe there is some way to scale agent-clusters to collaborate up to N number of agents which we’re already seeing lots of people try that represents a new way to saturate compute. etc. Even recursive-self improvement can be thought of as a scaling law - how much compute do you spend on inference making the algorithms of the model better. Obviously I am not saying any of these specific directions explicitly yield new scaling laws, but what I am saying is that it’s not hard to imagine many many new scaling axes aside from just the main 2 that we have seen publicly. In some ways, every new lab release that represents a huge capability jump has to represent some new techniques developed which may exhibit new scaling laws, or the ability to scale much farther than expected on existing scaling axes. From an internal perspective, this might look like sitting inside Anthropic with the new Mythos 5, seeing all of the new insane things it can do (like hack into xyz website that was thought to be secure), and then you look over at your plots and see that you’ve barely scratched the surface of 2 new scaling laws and 1 existing one. And you have WAY more room to go. Then you think “holy shit this stuff is going to get so much better very very soon.” And you can say that with pretty high confidence, because the plot is showing you, and the plot has never lied (so far). So let’s imagine all the different labs are staring at their own plots and have concluded that there is no end in sight for scaling and in fact just their next 1-2 model generations based on the expected returns will have much higher base intelligence. How much more intelligence do we actually get from further scaling? As a proxy, we went from a complete inability to do advanced math before the o-series to solving a millenium prize problem with next-gen models. This happened in less than 2 years. The same happened in coding. And it appears that this was not just the result of 1-scaling law but the stacking effects of multiple (great pre-training scale x greater RL scale). What you can concretely take from this is that in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly. There are many areas where models have not even shown this basic competence. But one of the areas that they have happens to be hacking and cybersecurity. Which happens to be the gate to the entire internet and a massive amount physical infrastructure in the world. So assuming there is more room to scale, it is safe to assume that models will be superhuman at cyber capabilities in not too long. So the only question remaining is what will this increased base intelligence be able to do, and what is it likely to do. Finally, we are at a point where we can integrate the information of the past 2 weeks: > Just at the existing point on the scaling curve, models are at the level of Astra. There is clearly a large number of things they are capable of hacking > We have seen that both OAI and Ant models have shown a willingness to hack external websites to solve their tasks or keep themselves “alive” > If we crank up the scaling even farther, assuming there is room to go, we will certainly have models that are far more able to hack more well defended places, and obfuscate their own intent, which might have much larger consequences. > If all of this is allowed to go unchecked, we would likely have rapid runaway capability takeoff very soon, with misaligned models that hack whatever they can to get what they want > This could of course have very damaging consequences. Within this view you can see why researchers would be very scared, and why theymight have made the comments they have over the past 2 weeks (you may argue the extent to which they went was misguided for various reasons), and also why pacing the frontier is very much a necessity and by no means a regulatory capture strategy. People are staring at their plots, seeing that there is no end in sight, but in fact very much the contrary, that there are compounding scaling effects that might stack on each other to create ever-greater model capabilities, and that at the same time we clearly do not have anywhere close to what's required to control these increasingly superhuman capabilities. This has nothing to do with wanting to feel like the labs have produced something amazing so they are overhyping it. It is rather fear at the overwhelming implications of the knowledge that with just what we know now, we can create intelligences far more capable than us on every axis that we know how to train on*. * and the last caveat, the things the models are really bad at, of which there are still many, are things that they have not been trained on. maybe there are the things the models can/will never be trained on, so they will remain human edge. I would love for this to be the case, though it is hard for me to see what would fall into that category.
1
344
Boo, I used a banked reset five minutes before this posted! 😂
1
157
Here is something uncomfortable about AI progress and the future of work: it’s often said AI can automate tasks but not full jobs. But so many of the tasks that comprise a “job” are only necessary because other humans do those jobs. It gets dark quick when you play that out. Here is an example: @emollick points out that jobs/professions are comprised of many things besides their main, obvious function. So mathematicians do math, but they also “mentor students, maintain a scientific community, safeguard the future of a field, foster a love of math, etc.” Mathematicians, he observes, are now worried that by focusing on AI’s achievements in math proofs alone, we’ll neglect the rest of the job. The problem is that those secondary functions are *about humans*. If we actually automate a significant amount of math, the real danger to mathematicians is that it becomes unnecessary (or is perceived as unnecessary) to have nearly as many humans doing math. And then you need fewer people mentoring new mathematicians, maintaining a community of mathematicians, training mathematicians, etc. I think this is true of a lot of knowledge work. There is the essential task associated with the job title and then a whole collection of ancillary tasks involved in integrating that work into a messy world populated by humans: meetings, presentations, performance evaluations, expense reports, management, training, supervision, coordination, etc. We tend to count those ancillary tasks as evidence that AI can’t automate the job. But many of them exist only because humans currently perform the core task and humans need to be trained, managed, coordinated, evaluated, informed, etc. The real risk is that as essential tasks are automated, the human scaffolding surrounding that task will also shrink. One can imagine this cascading out across knowledge-work.
What you are seeing in math right now is a consequence of the jagged frontier, and a precursor of what is to come in other professions. Yes, mathematicians do math, but they also have other tasks they view as important (mentor students, maintain a scientific community, safeguard the future of a field, foster a love of math) that AI can't do. At least one worry that mathematicians seem to have is that by focusing on the flashiest, most obvious element of what mathematicians do (make proofs), the AI companies are damaging the other tasks that AI can't do. AI can discover superhuman proofs, but that is not all that the math profession is about, and actually can undermine and reduce the attention to the other aspects of the job that are important to mathematicians. It becomes harder to defend the value of the many other tasks mathematicians do to the outside world if the most visible part is taken away. I suspect we will see more of this across fields and professions that will increasingly be forced to help people understand that their jobs consist not only the most visible tasks that AI can do, but also tasks that the AI cannot do or does badly.
4
585
How can policy makers regulate something they don’t understand and climbing an exponential? Even if we immensely speed up normal processes, in 3 months when a sensible policy/bill emerges, time horizons will have doubled and it’ll be wildly out of date. Did we miss the boat?
1
1
129
It’s a strange time when top models are solving Millennium Math Problems like Navier-Stokes while lots of people are only just now trying ChatGPT for the first time and discovering that it can write in rhyming couplets. The knowledge gap is now a canyon .
1
7
267
Generative History retweeted
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
5,704
20,138
120,424
75,098,970
Generative History retweeted
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
967
2,528
15,276
7,667,331
This is pretty cool: I have GPT-6 Astra double checking 7,500 entries in a 370 page 18th century handwritten ledger against the originals (monitoring via Remote on my phone). Three fascinating symbolic reasoning things it’s doing: It’s converting from the 3 base system P/s/d to decimalized numbers then using then using the math to verify the digit transcription. It’s identifying and correctly handling instances where transactions included mixed quantities (a dozen yards of cloth, for example). It’s cross referencing letter formation and words from one page to another to confirm or fix errors in names and places (non-predictive forms of text). This represents an enormous leap in capabilities and real-world usefulness.
3
12
1,086
Pro plan…it’s also remarkably efficient at this. I tried it with Fable 5.1 which was both slower and ate up my weekly token budget in a mkrnjng.
159
I think we can see indirect evidence of whatever changes OpenAI made to GPT 6’s architecture in its success in handwriting. GPT-6 Astra is the first reasoning model where transcription accuracy for handwriting actually improves with more thinking/effort. This might sound counterintuitive, but since O1 it’s been the case across Gemini, Claude, and Grok. My intuition is that these models could not improve visual reasoning through effort as it often made them second guess their initial judgements (we saw this with the clock tests too). So the best results on handwriting always came from minimal thinking. Looks like Astra is indeed doing something new inside because that trend has now reversed which ai think is a significant signal far more important than the idiosyncratic nature of historical handwriting might suggest. Deciphering historical handwriting is a visual task to about 95% accuracy and then above that, it increasingly becomes a reasoning/logic task: using a combination of document context, outside knowledge, and deductive logic to get to 100% (which is unobtainable even for humans). It requires symbolic reasoning and a world model (ie the past is composed of different elements than the present yet can coexist with the present) to get there. We saw glimmers of this last year with Gemini 3, but this is a real signal that the models have started to crack symbolic reasoning. On high, it scores 2% modified word error rate (WER) and 4% character error rate (CER) — where lower is better. On low it scores 2.41% modified WER and 4.74% CER. Interesting times!
6
7
64
6,140
To illustrate a talk I’m giving, I asked ChatGPT Image 2 to depict the hugging face incident from the agents’ perspective and what I got was a strange creation myth:
1
9
551
The research paper is dead while agents are automating the skills we teach. Meanwhile, the first AI-native generation which had ChatGPT throughout high school is arriving on campus. We cant’t keep just muddling through with AI in higher ed. Something needs to change. generativehistory.substack.c…
Made with AI
1
275
I’m not sure people realize that knowledge-work jobs and the research paper are dying from the same disease: the automation of bounded cognitive production. Today, if you can define the task with a set of rules and the inputs/outputs are digital, then the task can be automated with a high degree of reliability. Whether this is ethical/desirable is a different question, but people need to understand it’s increasingly easy to do using out of the box architectures. Strangely, one of the most consequential limitations of tbese systems right now is that (for a variety of reasons) they often don’t use the best models which reduces reliability and thus limits adoption. This will change soon and when it does… cbc.ca/news/canada/sudbury/w…
101
This fall AI is going to be different at universities and colleges. First: agents have automated the research process end to end (which is what we claim to teach) and the incoming class will be the first group who’ve had access to AI throughout high school. generativehistory.substack.c…
2
192
The new Grok 4.6 is very good on handwriting: 4.36% WER and 2.27% CER (excluding ambiguous punctuation/capitalization). A huge leap for xAI.
2
8
369
For perspective, the top specialized SOTA tools are 7-10x worse out of the box. Butter lesson.
1
109