3/ The fix: bootstrapping. Here's the idea in plain English:
• Take your N eval results
• Resample with replacement 1,000 times
• Compute your metric on each resample
• That spread IS your error bar
No distribution assumptions. Works on accuracy, cost, latency — anything.