System Design Concepts · API Design
Documentation: the docs come from the spec, not a hand-written copy
Hand-written docs are a second copy of what the code does, and copies go out of date. A spec that generates the reference, clients and checks cannot go out of date.
Two copies, and only one runs
- The reference page lists the same fields as the handler, but only the handler runs. The page was typed in by hand from the handler's code, and nothing links the two.
- Release v1.5 renames created to created_at, removes avatar_url and adds display_name in the handler. Nothing links the handler to the hand-written reference page, so the page does not change and is now wrong in three places.
- The pipeline runs on the v1.5 commit that made the page wrong, and all three checks pass. But none of the three checks reads the reference page, and no fourth check does, so the pipeline cannot see that the page is wrong.
- Nine days later, a partner writes their integration by reading the out-of-date reference page. So they send avatar_url, read created, and never learn that display_name exists.
- The partner's write gets a 400 for unknown field avatar_url, but reading created returns undefined without an error, so that code ships. The API team's bug costs the partner an afternoon, and the team finds out from a support ticket, not a failing build.
One spec, many outputs
- The spec is one OpenAPI file that describes the contract: 14 paths and 9 schemas, with 22 examples and 11 error codes. Both ways end at this one file: write the spec first and generate the code, or generate the spec from annotations in the code.
- The spec file generates every path, field and example in the browsable reference, its first output. So, unlike a hand-written page, the reference cannot go out of date as the code changes.
- The same file also generates client libraries in four languages, with their types. Nobody writes these clients by hand, so a field renamed in the spec changes in all four clients on the next release.
- A mock server answers requests with the examples from the spec, before the real endpoint exists. So the partner starts building on day one instead of waiting for the release.
- Now the file does more than describe the API, because the schema checks every request at the edge, where requests first arrive. The schema refuses any request with an undefined field, so that request comes back 400 and never reaches a handler.
- In CI, the same OpenAPI file is the fixture: the expected shape that contract tests compare every response against. A handler that silently drops a field fails the build, so CI would have caught chapter one's bug on the day it shipped.
- Generation from the spec stops here, so people still write four documents, and an integrator reads them first. The documents cover why to call this endpoint, how auth works end to end, how to handle each error, and migration after a breaking change.
Twenty minutes, or nine days
- A partner who opens the docs at minute zero must finish five steps, or gates, before one call works. On most APIs, the first gate is a form and a call back from sales.
- Gate one is a key: the partner signs up, creates a sandbox key on the page and pastes it, all in two minutes. A key that needs a sales call means three days of waiting before anyone writes a line of code.
- Gate two is a request that works as pasted, with real values and the partner's own key already in it, and the response printed beside it. This gate takes two minutes, compared with a day to build the same call from a written description.
- Gate three, their first failed call, matters most, and the API's error message is the most-read page in its docs. The 422 error message names the field, the rule it broke and what to send instead, so they fix the call in three minutes without a support ticket.
- Gate four is a sandbox with ready-made test accounts and fake money, so they can send twenty wrong requests before lunch. Without a sandbox, they test against production or wait for someone to make them an account.
- Gate five is what lets them commit to the API: published rate limits, a changelog and a deprecation policy with dates. Gate five also has signed webhooks and idempotency keys that stop a retry from charging twice, and twenty minutes in, they are live.
© LearnThatStack - diagrams may not be republished without permission.
The same API shape is written down twice, but only the copy in the handler code runs. A person typed the reference page from that code some time ago. Nothing links the reference page to the handler. Every documentation problem you have ever had starts with these two unlinked copies.
No line of code refers to the hand-written docs page. So release v1.5 changes only the handler. The handler renames the field created to created_at, removes avatar_url, and adds display_name. The docs page stays the same. So the page is now wrong in three places: one field renamed, one field removed, and one new field it never mentions.
The commit that made the docs page wrong passes the unit tests and the integration tests. The deploy then ships v1.5 on that same commit. There is no fourth check for the docs page, so zero checks read it. Nothing you own reads that page.
A partner has only the hand-written docs page, so the partner trusts it. Nine days later, the partner's integration sends avatar_url and reads the created field. The partner never learns that display_name exists. Now someone who does not work for you is doing your testing.
Two calls fail in two different ways, and the quiet failure is worse. The write gets a 400 response that names the field, so the bug gets fixed within an hour. The read returns undefined with no error, so the code that makes that call ships. Both failures reach you as a support ticket, the slowest way you have to find a bug.
There are two ways to get one OpenAPI file. Write the file first and generate the handlers from it, or annotate the handlers and generate the file from them. Both ways produce the same file: 14 paths and 9 schemas, with 22 examples and 11 error codes. Five outputs could come from that file, but today a person types every one of them.
The reference page is now an output of the spec file, so nobody types the page any more. Every path, field and example on the page comes from that file, so the page always matches the code. The difference from chapter one is in how the page is built. The page does not depend on anyone remembering to update it.
In chapter one, a renamed field broke the partner's integration. Here the same spec file generates client libraries in four languages. Nobody writes those libraries by hand. So a field renamed in the spec changes in all four libraries on the next release. A build step now handles the exact change that broke the partner's integration.
The examples in the spec file are enough for a mock server to answer requests. The mock server replies before the real endpoint exists. So your partner waits zero days and starts building on day one. The examples in a spec are also worth writing properly, because a server is going to return them.
At the edge, the schema refuses any request with a field the schema does not define. The schema returns a 400 that names that field. So the spec file has a new role: the document that describes the API now enforces the API. A document you can run cannot be wrong without anyone noticing.
Drift, where the code no longer matches the spec, now fails the build. In CI, contract tests use the spec file as the test fixture and check every response against the matching schema. So a handler that drops a field fails the build. Chapter one's bug would have failed the build on your machine the day the bug shipped, not nine days later in a stranger's ticket.
Generation from the spec stops here, and this is the honest half of the answer. People write why to call this endpoint and how auth works end to end. They also write what to do about each error, and migration notes for a breaking change. An integrator reads these four documents first. Generation removes drift, but it does not write the explanation.
Time to the first working call is a number you can measure and use as a design target. The measurement starts when a partner opens your docs. The partner then has to get past five gates, steps that block the call until each one is done. On most APIs, the first gate is a form and a call back from sales.
A key the partner can make on their own saves three days at the first gate, getting access. The partner signs up, creates a sandbox key on the page and pastes it in, all in two minutes. The key is scoped, so it cannot touch real money. A key that needs a sales call costs those three days before anyone writes a line of code.
An example that runs is worth more than a page that describes one. At gate two, the request runs exactly as the partner pasted it, with real values and the partner's own key already in it. The answer is printed beside the request, and the whole step takes two minutes. Building the same call from prose spread over four pages takes most of a day, and the call comes out wrong.
Gate three, the error response, is the one interviewers are listening for. The partner's first call fails with a 422 that carries a code the partner can look up. The error also names the field, the broken rule and what to send instead, so the fix takes three minutes, with no support ticket. The error you return is documentation, and integrators read it more closely than any other page.
Gate four is a sandbox, a safe place for the partner to be wrong. The sandbox holds seeded accounts (test accounts set up in advance) and fake money. So the partner can send twenty bad requests before lunch and has nothing to apologise for. Without a sandbox, the partner has two choices: test against production, or wait two days for someone to make them an account.
Gate five is not about the first call at all, but about turning one working call into a system somebody will build a business on. That system needs published rate limits, a changelog and a deprecation policy with dates. The same system also needs signed webhooks and idempotency keys, so a retry cannot charge twice. After twenty minutes, the partner is live.
A hand-written page is a second copy that nothing in your pipeline reads. So the page goes wrong without any warning. The fix is one OpenAPI file that generates the reference, client libraries and mock server. Request validation and contract tests come from the same file, so drift between code and spec fails a build. People still write the guides, error remedies and migration notes. For third-party integrations, name your design target: the time from opening your docs to one working call.
In interviews · 3 questions
Related Questions
- 01
- 02