I read an interesting study on token reuse and caching of content
Siddhant Khare (AI Researcher) turned some of these findings into an audit trail where an independent oracle predicts reusable tokens from token identity and cache policy, the engine attests, the prompt path has to show the same work, and cached and no-cache runs must emit identical token arrays.
Basically, zero reuse from cold state and zero reuse from eviction get separate verdicts, and timing stays null rather than zero, since token-level control cannot prove latency savings.
The study confirmed that reuse follows the first changed token, not request length. An interior edit after token three leaves three reusable tokens, a boundary edit leaves four, a changed suffix leaves eight.
This is very helpful to know, because we quote cache hit rates in roadmap reviews often, and now I am thinking we need to be smarter about what those are attesting or associated to