A clean benchmark result has an expiration date. Nobody prints it.
When a model is evaluated, the test set is new to it. That's the whole point. But public data doesn't stand still. The test set gets posted, cited, scraped, discussed, and folded into the next round of training corpora.
A year later, a newer model runs the same benchmark and scores higher. Some of that is real progress. Some of it is the test set leaking into training through the ordinary movement of the internet, with nobody doing anything wrong.
From the outside the two look identical. Same benchmark, higher number.
This is why an evaluation has to be bound to a dated snapshot of what the model could have seen. Without the date, a result isn't a measurement. It's a number whose meaning changes every time the corpus does.