Skip to content
APEXVYRA
Engineering

Reducing p95 latency: what actually moves the needle

RD

Ravi Deshmukh

Co-founder & CTO · Apr 18, 2026 · 6 min read

Email

We spent a quarter treating p95 latency as a single number to move, and mostly failed. It started moving once we stopped optimizing the median and looked at where the tail actually came from. Three changes to how APEXVYRA executes requests cut our own p95 by more than half — from around 1,850ms to under 820ms — without upgrading a single model. None of them are exotic. All three were boring enough that we almost didn't bother measuring them separately.

Why p50 numbers lie

A p50 chart that looks flat can hide a p95 that's climbing underneath it. Ours did exactly that: p50 held steady around 400ms for months while p95 crept from 1.2s to 1.85s. The tail was dominated by three things a median never sees — cold connections, one slow provider still sitting in the routing table, and requests that waited for a full response instead of streaming the first useful token. Fixing the median would not have touched any of them.

Keep-alive connections, not one per request

Every new HTTPS connection to a model provider costs a TLS handshake — 80 to 150ms depending on region, before a single token gets generated. Our worker pool opened a fresh connection per request because it was simpler to reason about. Switching to a shared, keep-alive connection pool, rotated every few minutes to avoid stale sockets, cut that setup cost to near zero for any request that wasn't the pool's first. That alone accounted for roughly 300ms of our p95 improvement.

python
import httpx

client = httpx.Client(
    limits=httpx.Limits(
        max_keepalive_connections=20,
        keepalive_expiry=120,
    ),
)

Speculative routing with a fast fallback

Some requests are latency-sensitive enough that waiting for the "right" model to respond costs more than getting a good-enough answer fast. For routes flagged latency-sensitive, we now fire the request to the Fast tier and the Balanced tier at the same time and return whichever responds first, provided it passes a lightweight guardrail check — the slower response is discarded. This doesn't reduce cost, since APEXVYRA bills for the response you keep, but it roughly doubles model-call volume for those specific requests. We scoped it to the ~15% of routes where users reported latency complaints, not every route by default.

Stream what you can

Waiting for a complete response before returning anything adds the model's entire generation time to your p95, even when the caller only needed the first sentence to start rendering. We stream tokens over Server-Sent Events wherever the caller supports it, and track "time to first token" as its own metric, separately from total completion time. Once we started measuring TTFT, it became obvious that some of our "slow" requests were actually fast responses that just weren't visible yet.

What we measured

MetricBeforeAfter
p50410ms395ms
p951,850ms810ms
p993,200ms1,640ms
Time to first token (streamed routes)—180ms

Where this doesn't help

None of this helps if the underlying problem is a provider having a bad day — a single degraded upstream model shows up in your tail regardless of how good your connection pooling is. Routing around a degraded provider, not just a slow one, is a separate problem we're still working through. Speculative routing isn't free either: budget for the extra model calls explicitly, and don't let it creep onto every route by default.

If you're chasing a p95 number, find out first which of these buckets it's actually coming from — connection overhead, provider variance, or full-response waiting — before reaching for a bigger model. In our case, the model was never the problem.

RD

Ravi Deshmukh

Co-founder & CTO at APEXVYRA. Spends most review cycles asking where the p99 went.

Comments

Discussion Demo — not persisted

  • SW

    Sam Whitfield 2 days ago

    The speculative routing point is underrated — we do something similar for our own internal search and it's the single biggest perceived-latency win we've shipped.

  • EK

    Elena Kowalski 1 day ago

    Curious what your keepalive_expiry tuning process looked like — did you land on 120s empirically or is that provider-recommended?