Because you don't do it all with an LLM.
This is why we need to move AI development beyond AI researchers who think about AI as a single engine that they are trying to get to do all the things.
That's not how this will work!
Why do I know this?
That's not how human brains even work.
Language processing is not the same part of the brain as reasoning. Counterfactual reasoning in the brain is like a network that reaches out to a bunch of different parts of the brain associated with memory, simulation and executive control.
Why would it be any different in AI? Why would you even try to process counterfactual reasoning in the same engine that processes language?
The answer is you don't!
This is not an LLM problem at this point. They have gotten good enough for their job and the extra tokens are not helping because they're just bogging the system down with more than it needs for the job it's supposed to do. This is a systems engineering and networking problem that can be solved with the combination of LLMs with the right memory, messaging and control centers.
Pioneer of causal AI, Judea Pearl, argues that no amount of scaling will get LLMs to AGI.
He believes current large language models face fundamental mathematical limitations that can't be solved by making them bigger.
"There are certain limitations, mathematical limitation that are not crossable by scaling up."
His core argument: LLMs don't learn how the world works. They learn from *human interpretations* of how the world works.
"What LLM's doing right now is they summarize world models authored by people like you and me available on the web and they do some sort of mysterious summary of it, rather than discovering those world models directly from the data."
He illustrates this with healthcare data.
When hospitals collect data on treatment effects, that raw data never reaches the LLMs.
Instead, the models consume doctors' written interpretations. Analyses shaped by people who already have a mental model of how disease and treatment work.
In other words, LLMs are learning from the map, not the territory.
The missing piece, according to Pearl, is causal reasoning — the ability to understand not just *what* happens, but *why*.
And he's clear this isn't a gap that more parameters or training data will close.
It raises a uncomfortable question...
If AGI requires machines that build their own world models from raw data rather than summarising ours, are we even on the right road?