
Dan Patiño
AI Strategy & Innovation at Coderhouse
Artificial Intelligence
What Is the KV Cache and Why LLM Inference Depends On It
The KV cache (key–value cache) is an inference optimization that stores the key and value tensors produced by each attention layer as an LLM generates tokens. During decoding, the model reuses those cached tensors instead of recomputing attention projections for every previous token on every new step. Without a KV cache, autoregressive generation would reprocess the entire context for each new token and become impractically slow as prompts and outputs grow.
Why it matters now: modern chatbots, agents, and RAG systems live or die on time-to-first-token and tokens-per-second. Serving stacks such as NVIDIA TensorRT-LLM treat the KV cache as first-class infrastructure, and technical explainers like How To Scale Your Model’s inference chapter show why memory for that cache often dominates long-context serving costs. If you already covered how models speed up speculative paths, our piece on speculative decoding pairs well with this concept: one technique cuts sequential decode latency; the other removes redundant attention work.
How Autoregressive Attention Creates the Need for a Cache
Decoder-only transformers predict one token at a time. At each step, attention compares the current token’s query vector against the keys of all previous tokens and mixes their values. The keys and values for past tokens do not change once written, so recomputing them every step wastes both FLOPs and memory bandwidth.
The KV cache stores those past keys and values per layer (and per head, depending on the attention variant). Prefill processes the full prompt in parallel and fills the cache. Decode then computes Q/K/V only for the newest token, appends its K/V, and attends against the growing cache to produce the next logits.
Prefill vs Decode
Prefill is compute-heavy and usually parallel across the prompt length: you materialize the full attention pattern once and save K/V. Decode is memory-bandwidth heavy: each step is tiny mathematically, but it must stream the entire cache from GPU memory. That is why long conversations feel “stuck” on hardware that has plenty of FLOPs but not enough memory bandwidth or KV capacity.
Why Memory Grows With Context
Cache size scales roughly with batch size × layers × sequence length × number of KV heads × head dimension × bytes per element × 2 (keys and values). Longer prompts, larger batches, and deeper models all inflate it. That linear growth is the reason context windows are not “free,” even when the model weights fit on the GPU.
Teams shrink the footprint with several complementary tricks:
Multi-query and grouped-query attention (MQA/GQA). Share K/V heads across many query heads so the cache stores fewer tensors.
Quantized KV caches. Store K/V in INT8 or INT4 to cut memory with careful quality checks.
Sliding windows and sparse attention. Keep only a recent window of K/V when full history is unnecessary.
PagedAttention and block reuse. Allocate the cache in non-contiguous blocks, share prefixes across requests, and reduce fragmentation in high-throughput servers.
How the KV Cache Connects to Product Latency
Product metrics map cleanly onto cache behavior. Time-to-first-token is dominated by prefill (and by how large the system prompt plus retrieved context is). Steady-state tokens-per-second is dominated by decode bandwidth against the cache. Concurrent users multiply batch pressure on KV memory, which is why inference platforms invest so heavily in eviction, offloading, and prefix reuse.
For AI engineers, that means prompt design is also systems design: a bloated system prompt, repeated tool schemas, or untrimmed RAG chunks burn KV memory before the model even starts answering. Pair retrieval quality work—see embeddings and RAG—with cache-aware serving so you do not retrieve more context than the GPU can hold.
Common Mistakes With the KV Cache
Assuming “bigger context window” equals better answers without measuring latency and cost. Forgetting that multi-turn chat appends to the same cache and can silently exhaust GPU memory. Disabling the cache in debugging builds and then wondering why generation is 10–100× slower. And treating speculative decoding, quantization, and KV paging as alternatives instead of layers of the same serving stack.
If you want to go deeper on building production data and AI systems, explore Coderhouse’s courses below.
AI Engineering Course — learn to build and serve LLMs with production inference patterns.
Data Analytics Course — develop the measurement skills to evaluate latency, cost, and quality trade-offs.
FAQ
Is the KV cache the same as a vector database cache?
No. A vector database caches or indexes embeddings for retrieval. The KV cache stores intermediate attention tensors inside the transformer during generation.
Does every LLM need a KV cache?
Any autoregressive transformer that attends over past tokens benefits from one. Some research architectures change attention, but production chat models almost universally rely on KV caching for decode.
Can the KV cache be shared across users?
Yes, when requests share a common prefix (system prompt, tool definitions, or a hot RAG chunk). Serving engines that support prefix reuse can keep one copy of those K/V blocks and attach divergent suffixes per user.
I'm Dan Patiño, head of AI Strategy & Innovation at Coderhouse. My day-to-day work involves merging the tactical management of e-commerce (CRO, Email Marketing and SEO) with the development of disruptive solutions. I specialize in building internal AI-powered apps to automate tasks and boost innovation within the team. I firmly believe that technology is strategy's best ally. To dive deeper into my professional journey, I'll be waiting for you on my LinkedIn profile.