ASSERT

What context will actually fit?

A model's memory is its weights plus its KV cache, and the cache is the half that decides how long a conversation you can serve. Get it wrong and vLLM does not warn you — it fails at startup, or it starts and dies on a long request.

The general formula is layers × kv_heads × head_dim × 2 × context. It assumes every layer caches, and that every layer caches over the whole context. This page does not assume either: it reads what each model declares about its own attention, and says which rule it applied.

Windows and hybrids: · ·
Plain attention: · ·

What the models declare, and what it changes

Everything below is read out of config.json. None of it is inferred from a model's name, and none of it is a correction factor — each is a different structure that caches a different amount.

Declared asWhat it means for the cacheEffect
layer_types Per-layer attention span. Gemma-4-12B has 40 of 48 layers on a 1024-token window and 8 on the full context; gpt-oss-20b is half and half at 128 tokens. 5.9× · 2.0×
layers_block_type Which layers attend at all. Nemotron-3.5 declares 52 layers of which 6 are attention; the rest keep a fixed-size recurrent state that does not grow with context. 8.7×
kv_lora_rank Latent attention caches one compressed vector per token per layer rather than K and V per head — so its own num_key_value_heads describes the attention, not the cache. 24.9×
sliding_window Only a window if the model says to use it. Some configs carry the field beside use_sliding_window: false, and applying it would understate the cache — promising a context that will not fit. —

The effect column is this model's cache at its own declared maximum context, against the general formula. Where a model really does cache every layer over the whole context — most of them — the two agree exactly, and the page says so rather than manufacturing a difference.