Skip to content
HandbookPublic
Security

Measuring performance where your users are

Lab scores and field data disagree for structural reasons. What to collect from real sessions, how to sample it, and how to hold a build to a budget.

Written forTeams whose users are on mid-range Android handsets and variable networks
Reading time9 min read
Last reviewed2026-09-02

A Lighthouse score is a simulation on a fixed device and network profile. It is repeatable and useful for catching regressions, and it is not what your users experience — particularly in India, where the device distribution has a long tail of three-to-five-year-old Android handsets and the network varies by an order of magnitude within a single session.

Why the two numbers disagree

Lab (Lighthouse)Field (real users)
DeviceOne simulated profileWhatever your users own, weighted by who visits most
NetworkA fixed throttleHandover between cells, congestion, a rickety office connection
Cache stateCold, every runMostly warm — repeat visitors dominate on an operations tool
InteractionNone, which is why INP cannot be measuredConstant, which is where INP problems live
SampleOne run, or fiveThousands, with a tail that matters more than the median

What to collect, and what to attach to it

A metric with no dimensions attached cannot be acted on: knowing your INP p75 is 480ms tells you nothing about which screen, which device class, or which release. Every measurement should arrive with enough context to be sliced.

ts
import { onLCP, onINP, onCLS, onTTFB } from 'web-vitals';

const context = () => ({
  route: routePattern(),                 // /work/[slug], never the resolved URL
  release: process.env.NEXT_PUBLIC_RELEASE,
  connection: (navigator as any).connection?.effectiveType ?? 'unknown',
  memory: (navigator as any).deviceMemory ?? null,   // proxy for device class
  saveData: (navigator as any).connection?.saveData ?? false,
});

const report = (metric: { name: string; value: number; rating: string }) =>
  navigator.sendBeacon('/rum', JSON.stringify({ ...metric, ...context() }));

[onLCP, onINP, onCLS, onTTFB].forEach((fn) => fn(report));
Collected on real sessions with the web-vitals library. sendBeacon survives the page being closed, which fetch does not.
  • Report the route pattern, not the URL. Per-URL data on a site with case-study slugs fragments into samples too small to read.
  • deviceMemory is a crude but effective device-class proxy. Segment by it and the "our site is fast" argument usually ends.
  • Keep the release identifier on every beacon. Without it you cannot answer whether last Tuesday made things worse.
  • Percentiles only. A mean over a long-tailed distribution is a number with no referent.

INP is where modern React sites lose

Interaction to Next Paint measures the worst interaction latency in a session — the tap that felt stuck. It cannot be measured in the lab, because the lab does not tap. On a heavy client-rendered page, the usual cause is a long task blocking the main thread while the handler waits its turn.

01
Find the long tasksPerformanceObserver on longtask, reported with the same context. Anything over 200ms during typical interaction is a candidate; over 500ms is the reason your INP is bad.
02
Yield to the browserBreak long synchronous work with scheduler.yield() where it is available, or a scheduled continuation where it is not. Rendering a thousand-row table in one pass is the classic case.
03
Render the response before the workPaint the state change immediately — the row moves, the button goes to a loading state — then do the expensive part. Perceived latency is what INP records.
04
Audit third-party scriptTag managers and chat widgets frequently own the worst long tasks on a page, and they are invisible in a lab run that blocks them. Load them after interaction where the business allows.

Enforcing a budget in CI

A performance budget that lives in a document is a wish. Put it in the build, fail on regression, and make the failure specific enough that the engineer who caused it knows what to do.

json
{
  "budgets": [
    { "path": "/*", "resourceSizes": [
        { "resourceType": "script", "budget": 180 },
        { "resourceType": "total",  "budget": 600 }
    ]},
    { "path": "/*", "timings": [
        { "metric": "largest-contentful-paint", "budget": 2200 },
        { "metric": "total-blocking-time",      "budget": 250 }
    ]}
  ]
}
Two budgets: one on shipped bytes, one on lab timings. The bytes budget catches more in practice.