Topic 1: Latency
11 min read·21 Sept 2026
By the end of this module, you'll be able to:
- Budget latency stage by stage, parallelize what can be parallelized, and think in p95 and p99
- Account for cost per answered question, and cut it with caching, prompt reuse, and model routing