A model's memory is its weights plus its KV cache, and the cache is the half that decides how long a conversation you can serve. Get it wrong and vLLM does not warn you — it fails at startup, or it starts and dies on a long request.
The general formula is layers × kv_heads × head_dim × 2 × context.
It assumes every layer caches, and that every layer caches over the whole
context. This page does not assume either: it reads what each model
declares about its own attention, and says which rule it applied.
Everything below is read out of config.json. None of it is
inferred from a model's name, and none of it is a correction factor — each is
a different structure that caches a different amount.
| Declared as | What it means for the cache | Effect |
|---|---|---|
layer_types |
Per-layer attention span. Gemma-4-12B has 40 of 48 layers on a 1024-token window and 8 on the full context; gpt-oss-20b is half and half at 128 tokens. | 5.9× · 2.0× |
layers_block_type |
Which layers attend at all. Nemotron-3.5 declares 52 layers of which 6 are attention; the rest keep a fixed-size recurrent state that does not grow with context. | 8.7× |
kv_lora_rank |
Latent attention caches one compressed vector per token per layer
rather than K and V per head — so its own
num_key_value_heads describes the attention, not the
cache. |
24.9× |
sliding_window |
Only a window if the model says to use it. Some configs carry the
field beside use_sliding_window: false, and applying it
would understate the cache — promising a context that will
not fit. |
— |
The effect column is this model's cache at its own declared maximum context, against the general formula. Where a model really does cache every layer over the whole context — most of them — the two agree exactly, and the page says so rather than manufacturing a difference.