Today we release my favorite episode of Training Data yet: the great Rich Sutton.
@RichardSSutton wrote the textbook, wrote The Bitter Lesson (and many other on-point essays like "Self-Verification, The Key to AI"), and trained a mafia of talented students who went on to change the AI landscape forever including David Silver, inventor of built AlphaGo.
@kjaved_ was Rich's PhD student at Alberta and wrote The Big World Hypothesis. They just left academia to start
@oaklab_ai
Their core argument:
(1) The Bitter Lesson: the world is massively more complex than any model of it, so anything trained on human-curated data has a ceiling
(2) Continual Learning: intelligence is continual by definition, and today's models stop learning the moment they ship.
The conversation covers:
— what The Bitter Lesson actually says, and what people get wrong
— why synthetic data is "just a big mistake," and the Big World Hypothesis behind it
— how LLMs are both a positive and a negative example of his own essay
— why no animal learns by supervised learning, and what squirrels can do that we can't
— the cure for catastrophic forgetting: per-weight step sizes and continual backprop
— why the biggest labs can't take a path where performance gets worse before it gets better
— a trillion parameters on 20 watts, and the Moore's Law math that makes it plausible
— why the endpoint isn't one mind but one design, running as many minds
It was both a fun generative idea- and debate-filled conversation, and a surprisingly human one too. Rich, thank you for beating cancer and changing the trajectory of AI. 💙
00:00 Introduction
02:10 An AI winter, a cancer diagnosis, and the move to Alberta
07:07 Writing "The Bitter Lesson," and what people get wrong
09:53 Are LLMs a positive or a negative example of it?
11:03 Synthetic data is "just a big mistake," and the Big World Hypothesis
18:01 AlphaGo, human priors, and why prior knowledge and learning should be friends
22:37 "Their weights never change": do LLM assistants actually learn?
26:09 Babies, squirrels, and why no animal learns by supervised learning
32:02 Rockets, imagination, and where paradigm shifts come from
36:42 The Alberta Plan and its 12 steps
38:53 Catastrophic forgetting and the cure
43:43 Oak's biggest ambition: a self-maintaining mind
47:56 Why the big labs are stuck in a local minimum
49:13 If everything goes right: LLMs, many minds, and hiring
The man who pioneered reinforcement learning thinks the rest of the field is weird, and lays it all out in today's episode. Together w/
@Alfred_Lin @sequoia