System Design Concepts · API Design
Authentication: proving who is calling
Authentication asks who you are and authorization asks what you may do. Authentication methods differ mainly in where the identity state lives.
A badge and the floors it opens
- A caller with no badge, no proof of identity, reaches the door, the identity check. The door cannot confirm who is calling, so the door refuses the request with 401: no identity yet.
- The caller signs in, so the door, the identity check, issues badge 124 with role staff. The lift, the permission check, reads that badge and opens floors 1 to 3 with 200.
- Floor 9 is pressed, and the lift refuses that floor with 403 while the identity check still shows badge 124 as valid. Identity passed at the door, but permission failed at the lift.
- The door is the first check, authentication, which asks who the caller is and fails with 401. The lift is the second check, authorization, which asks what the caller may do and fails with 403.
Where the session lives
- A client sends Cookie: sid=7f3a to the server. The server looks that id up in its session store and finds user 124. The server does that one lookup for every request.
- The client sends a JWT instead of a cookie. The server checks the signature and reads the claims from the token itself, sub 124 and a 60 minute expiry, so the server never reads the store.
- Three servers now run behind the load balancer. Each one checks the JWT with the same key and answers the request on its own. None of them reads the session store, so there is nothing to share between them.
- User 124 is banned, so the row for sid=7f3a is deleted from the store. The next request with that cookie looks up the id, finds no row, and gets 401.
- The ban is still in place, but the next JWT still has a valid signature, so server 3 answers 200 for a banned user. There is no row in the store to delete, so the token works until its 60 minutes have passed.
- Expiry drops from 60 to 15 minutes, and a refresh token, rt=9c1, gets its own row in the store. Deleting that row is the revocation, so a banned user's access token stops working within 15 minutes.
OAuth delegates without the password
- Three parties are involved: the app wants to read the user's files, the browser acts for the user, and the provider holds the account. The provider also keeps the password, and no login has happened yet.
- The wrong way. The app shows its own login form, and the user types the provider password into that form. The app now holds a copy of the password, so this path is ruled out.
- The app now makes and keeps a random code_verifier, v9q2, in place of the password copy, and sends only the verifier's hash, code_challenge 3e8f, with the redirect. The user allows read access to their files on the provider's own page.
- The provider issues a one-time code, a1b2, and redirects the browser back to the app with that code visible in the URL. The app keeps the code, beside the stored verifier.
- A copy of code a1b2, taken from the browser, tries the exchange on its own. The copy has no verifier, so the provider refuses the exchange.
- The app sends code a1b2 and verifier v9q2 to the provider on the back channel, a direct server-to-server call. The provider hashes the verifier and gets 3e8f, so the check passes and a token scoped to files.read goes back to the app. The password itself never travels anywhere.
- The app reads files with the token and gets 200. A write with the same token is refused with 403, because the token's scope is files.read and nothing more.
A key names an application
- The server creates a key for billing-worker and shows the key once. The app stores the key. The server stores only the key's sha256, 4c7a1d, in the api_keys row, next to the key's scope and a use count at 0.
- The api_keys table is copied out, and every row in the copy holds only a hash. A request that sends 4c7a1d as its key gets 401 from the server, which checked only the hash column.
- A GET request to /invoices with key k1 in the X-Api-Key header returns 200, so k1's use count rises to 1. The server refuses a POST to /invoices with the same key and returns 403, because the key's scope is invoices:read.
- A second key, k2, is created and shown once, so both keys work during the overlap. The app switches to k2, so k1's uses in the last hour drop to 0, and then the k1 row is marked retired.
© LearnThatStack - diagrams may not be republished without permission.
A request has to provide identity to access a resource from the server. Identity means proof of who is calling. This request has not provided any identity yet. So the server sends back the status code 401, Unauthorized.
The caller signs in. The server checks that identity and then issues badge 124 with the role staff, a proof of who the caller is. The badge allows access to floors 1 to 3, the resources this role may use, so the request gets a 200 OK. You now know who is calling, so the server can act on the caller's identity.
The server refuses the caller at Floor 9, a protected resource, and answers 403 Forbidden. Badge 124, the proof of who is calling, is still valid. Nothing about the badge changed, so a second sign-in would not help. The identity check passed, but this caller does not have the right permissions.
The two checks run in order. Authentication comes first, asks who you are, and returns 401 when that check fails. Authorization comes second, asks what you may do, and returns 403 when the caller lacks the right permissions. Usually the two checks are two pieces of middleware, in that order.
The identity badge, the proof of who is calling, still has to work on the next request. So something has to remember that badge, and the server keeps that memory in a session. The client sends Cookie: sid=7f3a, and the server looks that id up in its session store. The server finds user 124 there, so the request goes through.
The same client sends a JWT instead. The server checks the signature, then reads the claims straight from the token: sub 124 with a 60 minute expiry. The identity state is inside the token itself, so the server skips the store check entirely. A JWT is signed, not encrypted, so don't keep real secrets inside it.
Three servers now run behind the load balancer. Each server checks the JWT with the same key. Each server then reads the answers straight from the claims inside the token. No server reads the session store, so the servers have nothing to share. Adding a fourth server needs no shared session at all.
Now ban user 124. Delete the row for sid=7f3a from the session store. The next request with that cookie finds no row, so the request gets 401. Revocation with a session store is one delete, and the ban works on the very next request.
The ban is in place, and now a JWT arrives. Its signature is still valid, so server 3 answers 200 OK for a banned user. There is no session row anywhere to delete. So the token keeps working until its 60 minutes run out.
Drop the expiry from 60 minutes to 15, with a refresh token rt=9c1 in the store. Revocation now means deleting that refresh token row. A banned user's access token expires within 15 minutes instead of an hour. The window, the time a banned user can still call, gets shorter but does not close.
A JWT is not more secure than a session cookie. The token skips the store lookup but is slower to revoke, so you pick one based on the use case. Sessions work well for an app with a login and a logout button. Tokens work well in microservices, where many services can verify a token with the same key, so there is no shared session store and no lookup cost.
Sessions and tokens both work well when you hold the user's account. But a third app can also want the user's files. Then three parties take part: the app, the user, and the provider that holds the user's account.
The wrong way first: the app shows its own login form. The user types the provider password into that form. The app now has a copy of that password. So one breach at the app also breaches the provider account.
This time the app makes a random code_verifier, v9q2, and keeps that verifier. The app sends only the verifier's hash, the code_challenge 3e8f, with the redirect. The user goes to the provider's own page and allows read access to their files. So consent happens at the provider, the same place that stores the password.
The provider issues a one-time code, a1b2, and redirects the browser back to the app with that code. The redirect puts the code in the browser's URL, so the code is visible on the way. The app keeps the code for future requests. A code in a URL can be read from browser history or a log, so the code alone is not worth much.
A copy of code a1b2, taken from the URL, tries the exchange on its own. The copy has no verifier, so the provider refuses the exchange. The request stops at the door, the provider's check. PKCE, proof key for code exchange, makes a stolen code worthless.
The app sends the code a1b2 and the verifier v9q2 to the provider over the back channel. The provider hashes the verifier and gets 3e8f, so the check passes. The provider then returns a token to the app, scoped to files.read only. The password is never sent at any point in this exchange.
The app reads files with that token, and the provider returns 200 (OK). A write with the same token gets 403 instead. The provider refuses the write because the token's scope is files.read and nothing more. Two related flows are worth knowing: client credentials for service-to-service calls, and OIDC as the identity layer.
OAuth acts for a user. An API key is for an application instead, with no user involved. The app stores the key. The server stores only the key's sha256 hash, 4c7a1d, beside the key's scope and a use count of 0.
Someone copies the api_keys table out, but every row in that copy is a hash, not a key. A request that presents 4c7a1d, one of those hashes, as its key gets a 401 from the server. A leaked key table gives no access. That protection is the reason we store only the hash.
The server returns 200 for a GET to /invoices when key k1 is in the X-Api-Key header. The server counts that call, so k1's use count becomes 1. The server refuses a POST to /invoices with the same key and returns 403, because k1's scope is invoices:read, which only allows reading. Keys go in a header, not in the URL, because proxies and logs keep URLs.
Key rotation needs an overlap, so you create a second key, k2, and show it once. Both keys work for a while. The app switches to k2, so the number of k1 uses in the last hour drops to 0. Retire k1 then, because that count proves nothing uses k1 anymore.
So authentication asks who you are. Authorization asks what you may do. The methods differ mainly in where you keep the identity state. That state can be in your own store, in the token itself, or in a hashed key row.
In interviews · 4 questions
Related Questions
- 01
- 02
- 03
- 04