AI Inference & Serving

PagedAttention

An attention-cache management approach that maps logical token blocks to physical storage blocks instead of requiring each sequence's entire cache to occupy one contiguous allocation. The block mapping addresses allocation and sharing problems. It is not the same as paging a web document or discarding old conversation tokens.

Reviewed

Sources

Member lesson

The definition and sources are public. The complete practical lesson is for members.

Compare membership plans ยท Already a member? Sign in

Explore all dictionary definitions