Reducing p95 latency: what actually moves the needle
Ravi Deshmukh
Co-founder & CTO · Apr 18, 2026 · 6 min read
We spent a quarter treating p95 latency as a single number to move, and mostly failed. It started moving once we stopped optimizing the median and looked at where the tail actually came from. Three changes to how APEXVYRA executes requests cut our own p95 by more than half — from around 1,850ms to under 820ms — without upgrading a single model. None of them are exotic. All three were boring enough that we almost didn't bother measuring them separately.
Why p50 numbers lie
A p50 chart that looks flat can hide a p95 that's climbing underneath it. Ours did exactly that: p50 held steady around 400ms for months while p95 crept from 1.2s to 1.85s. The tail was dominated by three things a median never sees — cold connections, one slow provider still sitting in the routing table, and requests that waited for a full response instead of streaming the first useful token. Fixing the median would not have touched any of them.
Keep-alive connections, not one per request
Every new HTTPS connection to a model provider costs a TLS handshake — 80 to 150ms depending on region, before a single token gets generated. Our worker pool opened a fresh connection per request because it was simpler to reason about. Switching to a shared, keep-alive connection pool, rotated every few minutes to avoid stale sockets, cut that setup cost to near zero for any request that wasn't the pool's first. That alone accounted for roughly 300ms of our p95 improvement.
import httpx
client = httpx.Client(
limits=httpx.Limits(
max_keepalive_connections=20,
keepalive_expiry=120,
),
)
Speculative routing with a fast fallback
Some requests are latency-sensitive enough that waiting for the "right" model to respond costs more than getting a good-enough answer fast. For routes flagged latency-sensitive, we now fire the request to the Fast tier and the Balanced tier at the same time and return whichever responds first, provided it passes a lightweight guardrail check — the slower response is discarded. This doesn't reduce cost, since APEXVYRA bills for the response you keep, but it roughly doubles model-call volume for those specific requests. We scoped it to the ~15% of routes where users reported latency complaints, not every route by default.
Stream what you can
Waiting for a complete response before returning anything adds the model's entire generation time to your p95, even when the caller only needed the first sentence to start rendering. We stream tokens over Server-Sent Events wherever the caller supports it, and track "time to first token" as its own metric, separately from total completion time. Once we started measuring TTFT, it became obvious that some of our "slow" requests were actually fast responses that just weren't visible yet.
What we measured
| Metric | Before | After |
|---|---|---|
| p50 | 410ms | 395ms |
| p95 | 1,850ms | 810ms |
| p99 | 3,200ms | 1,640ms |
| Time to first token (streamed routes) | — | 180ms |
Where this doesn't help
None of this helps if the underlying problem is a provider having a bad day — a single degraded upstream model shows up in your tail regardless of how good your connection pooling is. Routing around a degraded provider, not just a slow one, is a separate problem we're still working through. Speculative routing isn't free either: budget for the extra model calls explicitly, and don't let it creep onto every route by default.
If you're chasing a p95 number, find out first which of these buckets it's actually coming from — connection overhead, provider variance, or full-response waiting — before reaching for a bigger model. In our case, the model was never the problem.