KV Cache Architecture
The KV cache is service state governing routing, isolation, reuse, and recovery—not merely a memory optimization.
Design block lifecycle, prefix reuse, tiered storage, eviction, and tenant boundaries around the workload.
01
Problem definition
As context and concurrency grow, cache fragmentation and eviction can bottleneck before compute.
Routing for reuse may worsen tenant isolation and tail latency.
- Long prefixes or system prompts repeat across requests
- Memory pressure and recomputation are material under long-context concurrency
- Requests are mostly short and non-repeating
- Cache keys and tenant isolation policy cannot be defined
KV Cache Architecture: system plate
- 01
State 1
Trace per-request cache occupancy, prefix duplication, and lifetime distributions.
- 02
State 2
Define an explicit lifecycle for block allocation, reference, replication, eviction, and recovery.
- 03
State 3
Measure bandwidth and recomputation break-even points across GPU, host, and remote tiers.
- 04
State 4
Validate isolation with failures including cache poisoning and cross-tenant leakage.
FEEDBACKFailed acceptance returns evidence to the first controlled stage: Prefix hit rate and recompute reduction.
- The workflow begins with Trace per-request cache occupancy, prefix duplication, and lifetime distributions..
- It reaches an acceptance decision through Prefix hit rate and recompute reduction.
03
Design method
The KV cache is service state governing routing, isolation, reuse, and recovery—not merely a memory optimization.
- 01
Stage 1
Trace per-request cache occupancy, prefix duplication, and lifetime distributions.
- 02
Stage 2
Define an explicit lifecycle for block allocation, reference, replication, eviction, and recovery.
- 03
Stage 3
Measure bandwidth and recomputation break-even points across GPU, host, and remote tiers.
- 04
Stage 4
Validate isolation with failures including cache poisoning and cross-tenant leakage.
04
Application scenarios
Long prefixes or system prompts repeat across requests
As context and concurrency grow, cache fragmentation and eviction can bottleneck before compute.
- APPROACH
- Trace per-request cache occupancy, prefix duplication, and lifetime distributions.
- BOUNDARY
- A higher hit rate does not always mean lower end-to-end latency.
Memory pressure and recomputation are material under long-context concurrency
Routing for reuse may worsen tenant isolation and tail latency.
- APPROACH
- Define an explicit lifecycle for block allocation, reference, replication, eviction, and recovery.
- BOUNDARY
- Sensitive prefixes prioritize isolation over reuse.
05
Design choices
| Decision | Gain | Cost | Watch |
|---|---|---|---|
| Long prefixes or system prompts repeat across requests | Trace per-request cache occupancy, prefix duplication, and lifetime distributions. | A higher hit rate does not always mean lower end-to-end latency. | Prefix hit rate and recompute reduction |
| Memory pressure and recomputation are material under long-context concurrency | Define an explicit lifecycle for block allocation, reference, replication, eviction, and recovery. | Sensitive prefixes prioritize isolation over reuse. | Cache occupancy, eviction, and fragmentation |
KV Cache Architecture: system plate
- As context and concurrency grow, cache fragmentation and eviction can bottleneck before compute.
- A higher hit rate does not always mean lower end-to-end latency.
- Define an explicit lifecycle for block allocation, reference, replication, eviction, and recovery.
- Prefix hit rate and recompute reduction
- Tail latency and isolation
- Sensitive prefixes prioritize isolation over reuse.
- The workflow begins with Trace per-request cache occupancy, prefix duplication, and lifetime distributions..
- It reaches an acceptance decision through Prefix hit rate and recompute reduction.
07
Validation plan
| Measure | Method | Pass condition | Caveat |
|---|---|---|---|
| Prefix hit rate and recompute reduction | Trace per-request cache occupancy, prefix duplication, and lifetime distributions. | Repeated runs satisfy the acceptance threshold agreed during discovery | A higher hit rate does not always mean lower end-to-end latency. |
| Cache occupancy, eviction, and fragmentation | Define an explicit lifecycle for block allocation, reference, replication, eviction, and recovery. | Repeated runs satisfy the acceptance threshold agreed during discovery | — |
| Tail latency and isolation | Measure bandwidth and recomputation break-even points across GPU, host, and remote tiers. | Repeated runs satisfy the acceptance threshold agreed during discovery | — |
08
Constraints and failure conditions
Requests are mostly short and non-repeating
A higher hit rate does not always mean lower end-to-end latency.
Cache keys and tenant isolation policy cannot be defined
Sensitive prefixes prioritize isolation over reuse.
09
Engagement model
- 01
Diagnosis
PattyAnalyze the current system and its failure signals.
ClientProvide representative work, data boundaries, and operating constraints.
Prefix hit rate and recompute reduction - 02
Design
PattyDefine an explicit lifecycle for block allocation, reference, replication, eviction, and recovery.
ClientConfirm owners and acceptance criteria.
Cache occupancy, eviction, and fragmentation - 03
Validation
PattyMeasure bandwidth and recomputation break-even points across GPU, host, and remote tiers.
ClientMake the production-transition or stop decision.
Tail latency and isolation
10
Durable deliverables
- KV Cache Architecture decision record
- The KV cache is service state governing routing, isolation, reuse, and recovery—not merely a memory optimization.Client-owned · Patty-reviewed
- Validation harness and acceptance criteria
- Prefix hit rate and recompute reduction · Cache occupancy, eviction, and fragmentation · Tail latency and isolationJointly maintained
- Operations and recovery runbook
- A higher hit rate does not always mean lower end-to-end latency. · Sensitive prefixes prioritize isolation over reuse.Operating-team owned
11
Terminology
- KV Cache Architecture
- Design block lifecycle, prefix reuse, tiered storage, eviction, and tenant boundaries around the workload.
- Acceptance criterion
- Prefix hit rate and recompute reduction
- Operating boundary
- A higher hit rate does not always mean lower end-to-end latency.
REFERENCES
References and primary material
- PagedAttention
Primary material for the method and terminology.
- vLLM Automatic Prefix Caching
Primary material for the method and terminology.