Skip to content
HandbookPublic
Security

Latency budgets that hold under load

How to decompose a response-time target into per-service budgets, find the work that actually costs you, and stop a slow dependency taking the system down with it.

Written forBackend engineers responsible for a system with a real traffic peak
Reading time11 min read
Last reviewed2026-08-18

Most API performance work starts after someone complains, which means it starts with an average. Averages hide the problem: a p50 of 120ms with a p99 of 4 seconds is a system where one request in a hundred is unusable, and on a screen that fires six requests, one user in six sees it. Decide the budget first, express it in percentiles, and hold the code to it.

Decompose the budget before optimising anything

A budget is only useful once it is split across the things that consume it. Write it down at design time and the architecture arguments resolve themselves — you stop debating whether a second service call is acceptable and start checking whether it fits.

SegmentBudget (p95)What eats it
Network in, India, 4G80 msTLS handshake on a cold connection; keep-alive matters more than compression
Gateway and auth15 msToken verification; cache the signing keys, never fetch per request
Application logic60 msSerialisation, validation, the work you actually meant to do
Database90 msUsually one query, occasionally forty — see the next section
Downstream calls150 msPayment gateways, SMS, WhatsApp; the segment you control least
Response out55 msPayload size, which is the part teams forget is latency
A 450ms p95 budget for an operations screen, as split on a real engagement. The numbers are ours; the method is the point.

The N+1 is still the most common cause

Every ORM makes it easy to issue one query per row without noticing, and the cost only appears when the list gets long — which is to say, in production, at the busiest hour. Do not rely on reading the code to find these; count queries per request in tests and fail the build.

python
@contextmanager
def max_queries(n: int):
    queries = []
    with connection.execute_wrapper(lambda ex, sql, p, m, ctx: (queries.append(sql), ex(sql, p, m, ctx))[1]):
        yield
    if len(queries) > n:
        raise AssertionError(
            f"{len(queries)} queries, budget {n}\n" + "\n".join(queries[:10])
        )

def test_pass_list_is_not_n_plus_one(client):
    create_passes(50)
    with max_queries(4):
        client.get("/api/passes?limit=50")
A query counter as a test fixture. This has caught more latency regressions for us than any profiler.

Payload shape is latency

A 900KB JSON response is not a bandwidth problem, it is a parse-and-render problem on the device holding it — and on an older Android handset at a gate, JSON parsing alone can cost more than the network transfer. Shape the response to the screen.

  • Return the fields the screen renders, not the whole row. A list endpoint and a detail endpoint want different projections, and sharing one serialiser between them is how payloads grow.
  • Paginate by cursor, not by offset. OFFSET 10000 makes the database count ten thousand rows it will discard, and the cost rises as the table grows.
  • Send dates as ISO strings and numbers as numbers. Client-side coercion of loosely typed payloads is a real, measurable cost at list scale.
  • Compress, but check: on very small responses the CPU cost of compression exceeds the transfer saving. Set a minimum size threshold rather than compressing everything.

Cache in layers, and name the invalidation

The question that decides a caching design is not where to put the cache — it is what makes the entry wrong. Write that down per layer before adding one, because a cache with no defined invalidation is a bug with a scheduled delivery date.

LayerGood forInvalidated by
CDN / edgePublic, identical-for-everyone responsesContent hash in the URL, or a purge on publish
Application (Redis)Expensive derived data — dashboards, aggregatesAn explicit key delete on the write path, never TTL alone
Per-request memoisationThe same lookup called five times in one requestThe end of the request; free correctness
Database materialised viewReporting queries over large tablesA scheduled refresh, with its staleness stated in the UI

Timeouts, retries and the failure you cause yourself

Most production outages we have been called into were not caused by the dependency that failed. They were caused by the caller's response to it: no timeout, so threads piled up; a retry with no backoff, so the recovering service was knocked over again; a retry on a non-idempotent write, so a payment was taken twice.

01
Every outbound call has a timeoutSet it from the budget, not from the library default — which is frequently infinite. If the segment budget is 150ms, a 30-second timeout is not a safety net, it is a queue.
02
Retry only what is safe to repeatReads and idempotent writes with an idempotency key. Anything that moves money or issues a credential retries only through a key the downstream service deduplicates on.
03
Back off, with jitterExponential backoff with full jitter. Fixed retry intervals synchronise every client in the fleet into a thundering herd at the exact moment the dependency is trying to recover.
04
Fail fast when the dependency is downA circuit breaker converts a slow failure into a fast one. A gate scan that fails in 200ms with "check the register" is operationally better than one that hangs for thirty seconds.
05
Shed load rather than collapseAbove a concurrency ceiling, reject with a clear error instead of accepting work you cannot complete. Partial service under peak beats total failure for everyone.

Measure the tail, in production

Load testing tells you what the system does under your assumptions. Production tells you what it does under everyone else's. Instrument for percentiles from the first sprint: p50, p95 and p99 per endpoint, plus queue depth and pool saturation, which are the two numbers that predict a collapse before latency shows it.