# AI KV Cache
An optimization that stores previously computed key and value tensors from the [[AI Attention|attention mechanism]] during autoregressive generation. Without it, the model would recompute attention over all previous tokens for every new token generated.
KV cache trades memory for speed: it makes generation O(n) instead of O(n^2) per token but grows linearly with sequence length. For long contexts, the KV cache can consume tens of gigabytes of GPU memory.
KV cache size is a primary constraint on long-context inference. A model might support 128K tokens in theory, but the KV cache memory required to actually serve that context at reasonable batch sizes can be prohibitive.
Techniques to manage it:
- **Paged attention** (vLLM): treats KV cache like virtual memory, eliminating fragmentation
- **Sliding window attention**: only caches the most recent N tokens
- **KV cache quantization**: storing keys/values in lower precision (e.g., FP8)
- **Multi-query / grouped-query attention (MQA/GQA)**: shares key-value heads across query heads, reducing cache size
- **Sparse attention with token-wise compression**: as in DeepSeek Sparse Attention (DSA); see [[DeepSeek v4]], which cuts KV cache size to ~10% of DeepSeek V3.2 at the same context length
Directly relevant to [[Context Window]] management and [[Context Compression]] strategies.
## Prefix caching as an API: Jev
A KV cache doesn't have to die with a single request. When many requests start with the same prefix (a system prompt, a long document), the keys and values for that prefix can be computed once and reused. vLLM's PagedAttention (Kwon et al., 2023) allows flexible sharing of the KV cache across requests. SGLang's RadixAttention (Zheng et al., 2023) goes further and reuses the KV cache automatically across requests with a common prefix. Provider-side prompt caching is the same idea, billed per token.
[[Jev]] (TypeSafe AI) turns that trick into the shape of the API. You send one state and many questions. Archer Hume reverse-engineered jev-1.13 in September 2026 with latency and token-accounting probes, and he concluded with high confidence that "every question reads the same state cache", with each branch adding only its own instructions and answer options. His math: with S state tokens and Q questions, separate requests process the state about Q times. Sharing brings that from Q×S down to S. TypeSafe hasn't published its architecture, so this is an outside inference, but it fits the measurements.
The practical consequence is that asking more questions is almost free. TypeSafe's parallel questions cookbook runs 13 questions over the GDPR Wikipedia article: one batched call cost $0.000497, 13 separate calls cost $0.006090. That's 12.2x cheaper (10x faster too, though that figure adds up serial latencies), with no change in the answers. The primitives page quotes slightly lower figures (11.5x and 9.6x). The saving approaches Nx as the document grows.
That's what makes [[Speculative Fan-Out]] rational: ask every question your code might need, including the ones that only matter for some inputs, and ignore the answers you don't use. Once the state is cached, the extra questions cost very little.
## References
- [Efficient Memory Management for Large Language Model Serving with PagedAttention (arXiv:2309.06180)](https://arxiv.org/abs/2309.06180)
- [SGLang: Efficient Execution of Structured Language Model Programs (arXiv:2312.07104)](https://arxiv.org/abs/2312.07104)
- [Jev's Architecture Unmasked (Archer Hume)](https://archerhume.com/posts/jevs-architecture-unmasked/)
- [Parallel questions cookbook (TypeSafe docs)](https://docs.typesafe.ai/cookbooks/parallel_questions)
## Related
- [[Large Language Models (LLMs)]]
- [[Context Window]]
- [[Context Compression]]
- [[AI Attention]]
- [[Transformers]]
- [[Jev]]
- [[Speculative Fan-Out]]