LLM inference
Mentioned in
- Upper Bounds for Local Model Throughput
- Write-up of Reiner Pope's Lecture: How GPT, Claude, and Gemini Are Actually Trained and Served
- Theoretical Upper Bounds for LLM Performance
- Open source AI needs cheap local speed
- Local models for real-time triage
- The localening is here
- Towards 1-click setup for local models in OpenClaw
- Agentic Engineering needs rigor, not just intuition