The Self-Deception of Evaluation
One of the hardest traps to avoid is self-deception. Often, we fall into its jaws without a fight and become a hostage who defends the kidnapper. And you can only see the cell after you're out of it.
Present day AI is optimized to be evaluated by humans. It impresses its adjudicators, and that masks which tasks it should be doing and still isn’t.
Goodhart's Law says that when a measure becomes a target, it stops being a good measure. If the focus is hitting a number, the behavior of the system changes in ways that compromise what the metric was supposed to measure.
AI is a wonderful technology and has the potential to change how the world functions. And although a lot has been promised, much less has been delivered as products people use reliably.
Processes that are done by many, but loved by few are perfect candidates to be handled by AI, yet it falls short in delivering solutions. For example: account reconciliation, project task prioritization, expense filing, fraud detection, and task assignment.
My point is: Are the current metrics used to evaluate AI good enough for the job? Or are we fooling ourselves with numbers that are optimized to impress us?