Cloud & Infrastructure
Cache Hit Rate
The fraction of measured requests or units served using a cache, with the unit and eligibility rule explicitly stated. Request-level hits and token-level reuse answer different questions. Neither alone describes correctness, latency improvement or the total cost of an AI workload.
Reviewed
Sources
Member lesson
The definition and sources are public. The complete practical lesson is for members.