we're using this in production. examples:
- eng team loves a specific PM's PRDs? extract their taste, generate PRDs that read the same
- AI gets the facts right but the output reads like slop. this fixes slop by anchoring to what "good" actually looks like
- tested it across PRDs, process docs, and other domains. adds quality back in automatically
the key insight: with AI its easy to get correct. the hard part is getting good. this is a system for bottling "good."
autoreason constructs taste from debate. but whose taste? the judges'. that's undirected optimization.
what if you only need one golden doc? anchor to it. extract criteria. hill-climb. when the eval saturates, the gap reveals criteria no human would have thought to specify. extract again. repeat.
debate uses a human-defined jury as the loss function. this RLs the loss function itself.