But with Jev’s default threshold, it looked 7 points worse.
At the default 0.5 cutoff, Jev scored 76%.
We tuned the threshold on human-labeled validation data, then tested it on held-out examples.
At 0.8, Jev reached 87%, with the same false-alarm and miss rates as Opus 5.