LearnThatStack Ace your next interview

System Design Concepts · API Design

HTTP caching: infrastructure you do not own, working for you

Declare freshness with Cache-Control and identity with ETag, and caches serve your traffic. The hard part is invalidation.

cache-controlFreshness keeps the request homethe fastest request never leavesBrowser Astored again/products60s/me45spage loads6requests sent3Browser B200 from CDNCDNstored again/products20sOriginpublic, max-age=60, s-maxage=20
cache-controlFreshness keeps the request homeBrowser Astored again/products60s/me45sloads6sent3Browser B200 from CDNCDNstored again/products20sOriginpublic, max-age=60, s-maxage=20

Freshness keeps the request home

  1. Browser A sends GET /products request to the origin. The origin replies 200 with Cache-Control: max-age=60. Browser's cache is still empty, so every request go all the way to the origin.
  2. The browser stores the reply in its cache under /products and starts a countdown of 60 seconds. The max-age value is a length of time: the server indicates this response remains valid for 60 seconds.
  3. The page loads 3 more times, and the browser's cache answers every GET /products request. No request is sent to the server, so the page loads counter rises to 4 and the requests sent counter stays at 1.
  4. A CDN now sits between the browsers and the server. Browser B sends its first GET /products request. The CDN has no copy, so it fetches from the server once. The server responds with public, max-age=60, s-maxage=20: public means a shared cache may keep a copy, and s-maxage sets how long the CDN's copy stays fresh. The CDN keeps the copy for 20 and answers browser B.
  5. Browser A sends a GET /me request, its own account page. The server responds with private, max-age=60. The CDN keeps no copy, because private means only that one user's browser may store the reply. Browser A stores the reply for 60. no-store would mean no cache stores the reply at all.
  6. Both /products copies reach 0 seconds and go stale. The next GET /products request leaves the browser, finds nothing usable at the CDN, and reaches the server. Requests sent to the server climbs to 3.
  7. The server responds 200 with public, max-age=60, s-maxage=20 again, and the CDN stores the reply on the way back, fresh for 20 seconds. Browser A stores its own copy for 60 seconds, so both copies are fresh again because every new reply restarts the freshness times.

© LearnThatStack - diagrams may not be republished without permission.

Browser A sends GET /products request, and the server responds 200 with Cache-Control: max-age=60. That header is a promise with a time limit: this answer remains true for 60 seconds. The server owns this header field as only the server knows how long its own data remains fresh. And only safe methods get the header in the response, so methods like POST don't get it.

The browser stores the reply in its cache, and a countdown starts at 60 seconds. The client now holds the copy and its countdown, and nothing on the server side records that this copy exists. The max-age value is seconds, counted from the moment the server send the response.

The next three page loads never reach the server, because the browser's own cache can answer these. Those three page loads sent zero requests. So four page loads have cost one request in total. A cache hit is the fastest response an API can give.

A CDN is now placed between browsers and the server, and browser B sends its first request. The CDN fetches from the server once, and because the answer is public, the CDN can keep a copy in a shared cache. The s-maxage=20 directive gives that shared cache its own freshness limit of 20 seconds, this is separate from the browser's 60 seconds. One request to the server can now serve a city, and the server only handles one request instead of thousands.

Browser A asks for /me, its own account page, and the server marks the response as private. The CDN passes the answer through and don't cache anything, because only that user's browser is allowed store a private answer. Anything personal has to be private, or a CDN hands one user's page to the next. The stricter option is no-store, which means never store the answer at all. This is different from no-cache rule which means always check first.

Both timers, the browser's and the CDN's, reach 0, so both cached copies are now stale. The next GET /products request leaves the browser again, misses at the CDN because the CDN's copy is stale too, and reaches the server.

The server reply again. The CDN and browser A both store the new reply, and restart their clocks: 20 seconds at the CDN, 60 seconds in the browser. And the new reply has full body for data that may not have changed.

The server replies 200 with a 48 KB body and ETag: "v7". The tag is the identity of this exact body: change one byte and the tag changes. A W/ prefix marks a weak tag, that means equivalent but not byte-identical.

The cache's timer drops to 0, so the copy is now stale. Stale does not mean gone: the client still holds the copy, including the tag. The client could reuse this, but now the client is not sure if its still fresh. So the client can ask a cheaper question instead of asking for the resource again: is this copy still good.

The request carries If-None-Match: "v7" to the server. So the question is no longer send me /products but has /products changed since "v7". The server compares the tag it has with the tag in the request, and they match. Cache-Control: no-cache means always revalidate before reusing the copy.

The server replies 304 Not Modified with no body. The round trip still happens, but the body is 0 bytes. A validator saves bandwidth, not latency. The cache marks its copy fresh again and the timer restarts, so a 304 also renews the cache.

This time the resource changed and the server's tag is "v8". The request's If-None-Match: "v7" no longer matches, so the server replies 200 with a full body and the new tag.

Last-Modified and If-Modified-Since are the coarser pair, and they work the same way. The server sends Last-Modified: Tue, 19 Aug 2026 10:41:07 GMT, the next request sends If-Modified-Since: 10:41:07, and the server replies with 304. The timestamp has one-second resolution and no content hash. One-second resolution is fine for files on disk, but not for anything that changes every second.

There are five places on the request path that can cache a response , from browser to the database. Each of these has their its own expiry time. The first three: the browser, the CDN and the reverse proxy can cache HTTP. The app cache, Redis in this example, is managed by your code.

A write request arrives at the server: PUT /price 25. The database updates price to 25, but nothing tells the four layers above it. Each of those four caches still holds the old price. We'll discuss strategies to shorten the time those copies remain stale, and each strategy has its some trade-offs.

The first strategy is to wait, and 60 seconds later the proxy's 30 second copy has expired, so the proxy fetches the new value, 25. But the CDN's copy still has 240 seconds left, so the CDN keeps answering 20, and stale reads increase. The answer is wrong for up to max-age. When the cached value is a price, every one of those reads can be a refund

The second strategy is to purge. A PURGE /price request is sent to the CDN, so the CDN drops its copy and stores the new value 25 in its place. A purge is an API call with a delay, and it works for a CDN but the same purge can't reach a copy inside a user's browser.

The third strategy is an event. The price update publishes price.changed, so the app cache refresh its value to new price 25. This is fast, but coupled, because the write path now knows the cache exists. The modern middle ground is stale-while-revalidate: the cache serves the old copy while fetching the new one.

The browser still has the stale copy, app.3f9c.js, but the deployed page points at app.a71e.js instead. So the request for the new name misses every cache layer and reaches the server, which returns the file marked as immutable. Nobody asks for the old copy again. Hashed asset URLs exist for this reason: do not change the copy, change the name.

Each of the four strategies has a cost. Wait leaves the cache wrong for up to max-age. Purge is separate work at every layer, and it can't reach the browser. An event on write couples the write path to the cache. Rename works, if you control the name.

Cache-Control sets how long a copy stays fresh and who can store it. An ETag lets a client ask whether the client's copy is still good. Invalidation is the effort: a TTL, a purge or an event, each with a cost. Rename when you control the name.

In interviews · 3 questions

Related Questions

Also helps with

← All concepts