Two rounds of fine-tuning on a decade of published judgment. Three hours of GPU time. A 70-probe gate built before a single training run. And the thing that actually won was one paragraph of standing positions in the system prompt, ten minutes to write, nothing to compute.
Fine-tuning captured voice and bounded behavior. It did not capture judgment.
The gate caught the tuned model fabricating a benchmark result nobody ran.
The research and the rulings are Keith Townsend's. The discussion is AI, built on his published lab record. Neither voice is his, and neither is pretending to be. Every figure traces back to the record.
The full lab, every number, and the raw detail:
labs.layer2c.com/labs/constr…
The podcast:
Apple Podcasts: podcasts.apple.com/us/podcas…
Spotify: open.spotify.com/show/03469R…
RSS: labs.layer2c.com/podcast.xml
YouTube: youtu.be/70_5bcbCCaI
Sep 3, 2026 · 6:48 PM UTC
5
588
