The stack under the site, tuned to the request.
Web servers, caching, TLS, and the application runtime behind them. We measure the request before changing anything, because the slow part is rarely where the team assumed it was, and a change that is not measured is a preference rather than an improvement.
Nobody optimises the part they cannot see.
A slow site gets attributed to whatever the team can observe. The developers see slow queries, so the database gets blamed. Someone runs a lab score, sees a low number for JavaScript, and a sprint disappears into bundle size. Meanwhile the resolver is taking ninety milliseconds before a single byte of the site has been requested, and no dashboard in the building shows it.
The only fix that survives contact with reality is to break the request into its phases and time each one, on real connections rather than on a laptop with a hundred megabit link. DNS, connection, TLS, server think time, transfer, and the blocking work in the browser. Once those are separate numbers, the argument about what to fix ends.
Usually the answer is unglamorous. The certificate chain is longer than it needs to be. Compression is on for HTML and off for the API responses. There is an N plus one query behind the one endpoint every page calls. A third-party tag is loaded synchronously in the head. None of these are hard to fix. They are just hard to find without measuring.
The same page, the same network, both ends measured.
Every bar is a phase of the request, and every reduction has a specific change behind it. Nothing here comes from a synthetic lab score.
- DNS−78ms
Resolver moved to anycast, TTLs sane
- TCP + TLS−146ms
TLS 1.3, session resumption, OCSP stapled
- Server think−264ms
N+1 queries removed, pooling fixed
- Transfer−107ms
Brotli, correct cache headers
- Blocking JS−204ms
Third-party tags deferred
Measured on the same page, same network profile, before and after. Nothing here comes from a synthetic lab score.
What we take on
Web server and reverse proxy
One configuration that is readable, versioned, and identical across environments.
nginx, Apache, or Caddy configured from a template rather than by accretion, with a proxy layer that handles TLS, compression, timeouts, and rate limiting in one place. Staging serves the same config as production, so a difference in behaviour is a bug rather than an explanation.
TLS and certificates
Automated issuance, a short chain, and renewals that are monitored rather than trusted.
Modern protocol versions only, session resumption and OCSP stapling on, and issuance automated end to end. Because renewal automation fails silently, the certificates are then watched by the platform, so a renewal job that stopped running shows up weeks before anything expires.
Caching and delivery
Correct cache headers first, then a CDN in front of something worth caching.
Cacheability decided per route rather than globally, with headers that let the browser, the proxy, and the edge each do their part. Static assets fingerprinted and served immutable, HTML cached with revalidation where the application allows it, and a purge path that works during a deploy.
Application runtime
Process managers, workers, and pools sized against the traffic they actually get.
PHP-FPM pools, Node processes, or Python workers sized from concurrency measurements instead of defaults, with timeouts set at every layer so a slow dependency sheds load rather than queueing until the box falls over. Connection pooling to the database included, because that is where a surprising amount of the think time lives.
Deployment
A deploy that is repeatable, reversible, and does not drop requests.
Build once, promote the same artefact, drain connections before a restart, and keep the previous release ready to roll back to. Migrations separated from deploys so a schema change is a decision rather than a side effect of pushing to a branch.
Resilience under load
Rate limits, timeouts, and an origin that stays up when traffic arrives unexpectedly.
Load tested to find the breaking point before it finds you, with rate limiting at the edge, sensible connection limits at the origin, and a static fallback for the moment the application tier is genuinely gone. Pairs with the platform's DDoS protection where the traffic is hostile rather than merely heavy.
Measure, diagnose, change, measure again.
Measure
Real user timings and synthetic runs from the regions your traffic comes from, broken into phases, with the current numbers written down.
Diagnose
A report ordering the findings by milliseconds available, not by how interesting they are, with the effort each one takes next to it.
Change
Fixes applied in order, one at a time, in staging and then production, so each change owns its share of the improvement.
Re-measure
The same measurements repeated on the same profile, and the before and after handed over with the configuration.
Get the measurement first.
A week of timing your real traffic usually settles the argument about what to fix, and it costs less than the sprint spent guessing.