Some code review tools score better when nobody interacts with the review. Others score better when humans are actively reading and responding.
The difference isn't quality. It's how each tool is being used.
The benchmark now separates the two so you can compare fairly.
The software factory is already here.
We're seeing bots write code, bots review it, and humans reduced to dispatching the next tool in the chain.
Using 500k+ PRs from Code Review Bench, we looked at one question: can the human leave the loop yet?