Latency budgets that hold under load
How to decompose a response-time target into per-service budgets, find the work that actually costs you, and stop a slow dependency taking the system down with it.
Most API performance work starts after someone complains, which means it starts with an average. Averages hide the problem: a p50 of 120ms with a p99 of 4 seconds is a system where one request in a hundred is unusable, and on a screen that fires six requests, one user in six sees it. Decide the budget first, express it in percentiles, and hold the code to it.
Decompose the budget before optimising anything
A budget is only useful once it is split across the things that consume it. Write it down at design time and the architecture arguments resolve themselves — you stop debating whether a second service call is acceptable and start checking whether it fits.
| Segment | Budget (p95) | What eats it |
|---|---|---|
| Network in, India, 4G | 80 ms | TLS handshake on a cold connection; keep-alive matters more than compression |
| Gateway and auth | 15 ms | Token verification; cache the signing keys, never fetch per request |
| Application logic | 60 ms | Serialisation, validation, the work you actually meant to do |
| Database | 90 ms | Usually one query, occasionally forty — see the next section |
| Downstream calls | 150 ms | Payment gateways, SMS, WhatsApp; the segment you control least |
| Response out | 55 ms | Payload size, which is the part teams forget is latency |
The N+1 is still the most common cause
Every ORM makes it easy to issue one query per row without noticing, and the cost only appears when the list gets long — which is to say, in production, at the busiest hour. Do not rely on reading the code to find these; count queries per request in tests and fail the build.
@contextmanager
def max_queries(n: int):
queries = []
with connection.execute_wrapper(lambda ex, sql, p, m, ctx: (queries.append(sql), ex(sql, p, m, ctx))[1]):
yield
if len(queries) > n:
raise AssertionError(
f"{len(queries)} queries, budget {n}\n" + "\n".join(queries[:10])
)
def test_pass_list_is_not_n_plus_one(client):
create_passes(50)
with max_queries(4):
client.get("/api/passes?limit=50")Payload shape is latency
A 900KB JSON response is not a bandwidth problem, it is a parse-and-render problem on the device holding it — and on an older Android handset at a gate, JSON parsing alone can cost more than the network transfer. Shape the response to the screen.
- Return the fields the screen renders, not the whole row. A list endpoint and a detail endpoint want different projections, and sharing one serialiser between them is how payloads grow.
- Paginate by cursor, not by offset. OFFSET 10000 makes the database count ten thousand rows it will discard, and the cost rises as the table grows.
- Send dates as ISO strings and numbers as numbers. Client-side coercion of loosely typed payloads is a real, measurable cost at list scale.
- Compress, but check: on very small responses the CPU cost of compression exceeds the transfer saving. Set a minimum size threshold rather than compressing everything.
Cache in layers, and name the invalidation
The question that decides a caching design is not where to put the cache — it is what makes the entry wrong. Write that down per layer before adding one, because a cache with no defined invalidation is a bug with a scheduled delivery date.
| Layer | Good for | Invalidated by |
|---|---|---|
| CDN / edge | Public, identical-for-everyone responses | Content hash in the URL, or a purge on publish |
| Application (Redis) | Expensive derived data — dashboards, aggregates | An explicit key delete on the write path, never TTL alone |
| Per-request memoisation | The same lookup called five times in one request | The end of the request; free correctness |
| Database materialised view | Reporting queries over large tables | A scheduled refresh, with its staleness stated in the UI |
Timeouts, retries and the failure you cause yourself
Most production outages we have been called into were not caused by the dependency that failed. They were caused by the caller's response to it: no timeout, so threads piled up; a retry with no backoff, so the recovering service was knocked over again; a retry on a non-idempotent write, so a payment was taken twice.
Measure the tail, in production
Load testing tells you what the system does under your assumptions. Production tells you what it does under everyone else's. Instrument for percentiles from the first sprint: p50, p95 and p99 per endpoint, plus queue depth and pool saturation, which are the two numbers that predict a collapse before latency shows it.