System Design Concepts · API Design
Who controls the response shape
In REST the server decides the response shape, in GraphQL the client builds it, and in gRPC a compiled contract sets it. The trade-offs come from who decides the shape.
Over-fetching and under-fetching
- A profile screen on a phone needs the avatar, the name, and the last three orders, each with a product name and a total. The client has not sent a request yet, so all four rows on the screen are still waiting for data.
- The server answers GET /users/123 with 24 fields and 2400 bytes, but the screen uses only two: the name and the avatar. The other 22 fields still cross the phone's mobile radio, but the screen draws none of them.
- The reply from GET /users/123 held order_ids, not orders, so the client has to send a second request. GET /orders returns 54 fields and 3300 bytes, so the three order rows get their totals, but each product is still only an id.
- A third request turns the product ids into names, and the screen draws at 610 ms, using 8 of the 99 fields that came back. The three trips moved 6900 bytes and ran one after another, because each later trip needed ids from the one before.
- REST answers with an endpoint shaped like one screen: GET /screens/profile returns the 8 fields the screen draws, 600 bytes in one trip at 250 ms. But the order screen then needs its own endpoint, and so does the next screen.
- One endpoint per screen works, but every new screen needs a change on the server. The question behind all of this is who decides the response shape, and REST's answer is the server, one endpoint at a time.
The client builds the query
- The three REST paths are now behind POST /graphql, the only endpoint. Behind that endpoint is a schema that lists every type, every field, and what each field returns.
- The client writes the shape it wants as a query, one field per line. Every field already exists in the schema, the server's list of types and fields, but the server chose none of these fields.
- One response comes back in the shape of the query: 1 trip, 600 bytes and 250 ms, the same three numbers as chapter 1's purpose-built endpoint. So the difference is not speed, but that the client changed the shape and nobody deployed anything.
- In chapter 1 every GET was a cache key, so a CDN could answer the request without reaching the server. GraphQL sends every query to one URL with one method, POST /graphql, inside a body that a cache does not read. So HTTP caching is off, and a GET transport with persisted queries gets some of that caching back.
- The server resolves the product field with one lookup per order, so one HTTP request for three orders costs five database queries. The flexible query means the server now owns a query planner: a batching loader combines the product lookups into one and cuts the count to three.
- A client-built query can crash the database, and this query passes the depth limit of 8 because it nests five fields. But the query costs a million nodes against a cost limit of 1000, so the server's guard refuses it with a 400 before any resolver runs.
The contract compiles the call
- In gRPC the contract is a file in the repository beside the code, orders.proto, which names one service, one method and the response message. The number after each field is what goes over the network, so a field's number can never change.
- A compiler writes the client and server stubs from orders.proto, so nobody types them by hand and they cannot get out of sync. Renaming product_name in orders.proto makes both sides stop compiling, so the mistake fails the build instead of becoming a 400 in production.
- The same three fields go out as 22 bytes instead of 61, because protobuf, gRPC's binary format, writes field numbers, not field names. HTTP/2 then sends four calls over one connection instead of four separate connections, so no call waits behind another.
- The contract also sets how many messages are sent, and three of its four call shapes keep the stream open: server streaming, client streaming, and both at once. A REST endpoint gives only unary, one message out and one back, so the other three have to be built by hand.
- A browser cannot make the same call, because its fetch API cannot set the HTTP/2 frames that gRPC needs. So a grpc-web proxy has to sit in the middle and translate, and the network tab shows 22 bytes of binary, not a readable body.
- A real system runs all three at once, with REST at the public edge because anything can call it and a CDN can cache it. GraphQL runs where many screens want their own shapes, and gRPC runs between internal services with a compiled contract and small calls.
© LearnThatStack - diagrams may not be republished without permission.
A profile screen needs its data in one exact shape. The screen has four rows: the name with its avatar, then the last three orders, each with a product name and a total. But the server that answers the request has never seen this screen. Every problem in this concept comes from the server not knowing that shape.
The first request shows over-fetching: the response is bigger than the screen needs. The call to GET /users/123 returns 24 fields in 2400 bytes, but the screen shows only two of them. The mobile radio still turns on to download the other 22 fields, which the screen never shows. The endpoint is not badly written: it is written once for every caller, and this screen is only one of those callers.
Under-fetching is the same mistake in the opposite direction. This time the fixed response shape is too small, because the user record holds order_ids, not the orders themselves. So the client has to send a second request. The second request cannot start until the first one has finished.
On a phone, 610 ms means most of a second of blank screen. The screen waits for a third request to turn the product ids into names. The three requests could not run at the same time, because each request needed ids from the one before it. The three requests also moved 6900 bytes so the screen could show 8 of the 99 fields that came back.
REST has a reasonable fix: GET /screens/profile, an endpoint that returns exactly the eight fields the profile screen shows. So the profile call takes one trip, 600 bytes and 250 ms. But the order screen and the settings screen each need their own endpoint, so every new screen means a deploy. A ?fields= parameter on the old endpoint does the same job with less setup, but the server still decides which fields exist.
Speed was never the point, because the purpose-built endpoint was as fast as anything that comes after it. The real question is who decides the response shape. In REST, the server decides the shape, one endpoint at a time. The next two chapters give the other two answers.
GraphQL starts by putting the three REST paths behind one endpoint, POST /graphql. A schema behind that endpoint names every type and every field. The three REST paths still exist, and most GraphQL layers run in front of services that are still REST-shaped. So GraphQL does not replace REST.
The main idea of GraphQL is that the client writes the response shape it wants. Each line of the query names a field that already exists in the schema. So the server can check the whole query before running any part of it. In this query, the server decided nothing.
GraphQL is no faster than the purpose-built endpoint from chapter 1. GraphQL and that endpoint both cut three round trips to one, and 6900 bytes to 600. Both also cut the time from 610 ms to 250 ms. What changed is who decides the response shape. This time the client changed the shape, so nobody deployed anything.
The cost is HTTP caching. In chapter 1, every GET was a cache key, so a CDN or browser could answer that GET without calling your servers. Now every query is a POST to one URL, /graphql, with the query in the request body, which a cache does not read. A GET transport with persisted queries gets some of that caching back, but this setup adds a build step.
One HTTP request is not one database query, because the product field resolves once for every order. So three orders cost five queries, and fifty orders on a list screen would cost fifty-two. A batching loader collects the product lookups into one query, which cuts this N+1 problem to three queries. A flexible query means you now own a query planner.
A depth limit alone was never enough. This query nests five fields, inside the depth limit of 8, but it costs a million nodes against a cost limit of 1000. The guard on the endpoint refuses the query with a 400 before any resolver runs, so the query never reaches your database. A client can build a query that crashes your database, so you need a cost budget too.
In gRPC, a file checked into the repository beside the code decides the response shape. That file, orders.proto, names one service, one method, and the message the server returns. The number after each field in the message is what gets sent over the network. So once a field has the number 2, that field must keep the number 2 forever.
In gRPC, a mismatch between client and server is now a compile error. The compiler writes the client and server stubs from one file, so nobody types them by hand and they cannot get out of sync. Rename product_name in the file, and both projects stop compiling. A REST client would have found the same mistake in production, when a user reported a 400 error.
gRPC is fast for two reasons, and both matter. First, gRPC encodes messages with protobuf, a binary format that writes the field number 2 where JSON writes the text total_cents. So the same three fields take 22 bytes instead of 61. Second, HTTP/2 sends the four calls over one connection instead of four separate connections, so no call waits in a queue behind another call.
The gRPC contract file covers more than the fields. It also decides how many messages each call sends. A REST endpoint gives you only one of the four call shapes, called unary: one message out and one back. The other three keep the stream open: one out and many back, many out and one back, or both directions at once. With REST, you build those three yourself.
gRPC has a cost in the browser, because the browser cannot make a gRPC call on its own. The browser's fetch API cannot set the HTTP/2 frames that gRPC needs. So a grpc-web proxy between the browser and the server has to translate each call. The browser's network tab shows 22 bytes of binary, not a body you can read. This cost is why gRPC is used inside a system.
No single choice of who decides the shape fits every caller, so a real system runs REST, GraphQL and gRPC together. REST runs at the public edge, because any client can call a REST endpoint and a CDN can cache REST responses. GraphQL runs where many screens each want their own response shape, and you would rather not deploy again for each screen. gRPC runs between your own services, where both ends are your code and the contract compiles.
Answer with who controls the response shape, not with a feature list. In REST the server controls the shape of each endpoint, so you get fixed shapes and extra round trips. In GraphQL the client controls the shape of each query, so you lose HTTP caching and must handle the N+1 problem and query cost. In gRPC a contract compiled into both client and server controls the shape, so you lose direct browser calls and readable messages. Finish by saying that most real systems run all three, and where you would put each one.
In interviews · 4 questions
Related Questions
- 01
- 02
- 03