Mechanism◆Article
Gemma 4's E2B and E4B models share KV cache tensors across layers, cutting KV cache memory by roughly half in long-context settings.
Gemma 4's small models reuse KV projections from earlier layers instead of recomputing them in every layer, roughly halving KV cache size and saving several GB of memory at 128K context. ✦ AI generated
Sebastian Raschka · Ahead of AI · 2026-05-16 · original ↗
Since we share roughly half of the KVs across layers, we save approximately half of the KV cache size. For the smallest E2B model, this results in a 2.7 GB saving (at bfloat16 precision) in long 128K contexts, as shown below.
Read full article ↗excerpt · fair-use quotation
- ·E2B/E4B share KV tensors across layers
- ·Roughly half of KVs reused, not recomputed
- ·Cuts KV cache size by ~50%
- ·Saves 2.7 GB for E2B at 128K context
- ·Some layers reuse earlier layers' KV projections
- ·Saving ~half of KV cache size overall
- ·2.7 GB saved at bfloat16, 128K context
- ·Effect grows with longer context lengths
Around this claim