This is ideal to Run when ordinary success metrics stay healthy but interaction traces show the system is winning the task while the user is losing the thread. The clearest triggers are behavioral, not model-quality scores.
Watch for repeated corrections that do not stick: the user restates a constraint, tone, audience, or “don’t do X,” and a later turn drops it.
High edit distance after an accepted answer, frequent undo or regenerate on the same request, and sessions that end with the user rewriting the output themselves are stronger signals than a low star rating.
So are abandoned flows after an apparently complete answer, short follow-ups such as “no, I meant…,” and users narrowing scope mid-task because the system expanded it.
Product metrics that should prompt a sample are a rising gap between task-completion or thumbs-up rates and retention, reuse, or downstream acceptance of the artifact.
Look also for more tool actions than the user asked for, memory or profile fields changing without an explicit update, and recommendations or drafts that converge on one style even when the user keeps supplying counterexamples.
Support tickets and session notes that mention “it ignored me,” “it decided for me,” or “that’s not how I work” are direct reasons to score those transcripts.
It is less urgent when failures are simple factual errors, refusals, or latency.
Those already show up in standard evals.
This can be used for the cases where the answer looks finished and the user still had to fight the system to keep their goal, constraints, control, or self-representation intact.