Even the best models train on a lossy residue of the world. Language is not reality; it is a slow social compression of it. Spoken English carries tens of bits per second of new information, well below a 56k modem. Text is denser but still a thin pipe. Almost everything a nervous system registered never entered the corpus.
Give a model the uncompressed stream—electromagnetic waves and force, not tokens—and the first problem is no longer prediction but what to treat as a unit.
Tokens are an accident of human writing. A mind that saw the raw field would invent its own compression: events, invariants, causal structure, geometries we do not yet name.