Inference throughput Knowledge graph concept · also referred to as tok/s, serving throughput, useful serving throughput, throughput ceiling, tokens per second, token throughput Mentioned in Report Per-Session and Aggregate Token Throughput X thread · 2026-07-21 Upper Bounds for Local Model Throughput X thread · 2026-07-17 Open source AI needs cheap local speed X thread · 2026-07-01 Related concepts Inference scaling 2 KV cache 2 LLM inference 2 Speculative decoding 2 AI infrastructure 1 Local model deployment 1 Memory bandwidth 1 Memory capacity 1