request hedging
Request hedging is not a retry: how to tame P99 latency
Retries respond to failures, hedges respond to time. Firing a duplicate request after a short delay and taking the first response can collapse P99 latency at a small steady cost, and in 2026 it has become a production norm for LLM inference.