KV cache Knowledge graph concept · also referred to as KV fetches, per-conversation state, context traffic, KV, per-session context traffic, kv cache size per session Mentioned in Report Per-Session and Aggregate Token Throughput X thread · 2026-07-21 Upper Bounds for Local Model Throughput X thread · 2026-07-17 Write-up of Reiner Pope's Lecture: How GPT, Claude, and Gemini Are Actually Trained and Served post · 2026-07-14 Theoretical Upper Bounds for LLM Performance post · 2026-07-14 Related concepts LLM inference 3 Memory bandwidth 3 Speculative decoding 3 Inference batching 2 Inference scaling 2 Inference throughput 2 Memory capacity 2 Model quantization 2