Docs As Code
FeaturesLong read

What Good Code Documentation Actually Looks Like

Documentation that explains why code exists matters far more than restating what it does.

Editor-at-Large · · 10 min read
Cover illustration for “What Good Code Documentation Actually Looks Like”
Features · October 7, 2026 · 10 min read · 2,178 words

Most codebases today are not short on documentation.

The structural cause is familiar to anyone who has shipped software under a deadline.

A newer cause has joined the older one, and together they change the shape of the problem, not just its scale. AI tools now generate documentation quickly and with total confidence, but the 2026 version of the problem has shifted from quantity to accuracy. The result is a wave of polished, fluent, plausible-sounding documentation that carries no genuine intent behind it, confident in tone and unreliable on exactly the decisions a reader most needs explained.

What documentation is for

Good documentation exists to explain why code does what it does. The code itself already explains the what: a well-written function, read line by line, already shows its own mechanics to anyone fluent in the language. If documentation simply re-describes those mechanics, it is wasted effort dressed up as diligence.

What earns its place is the non-obvious: that a function caches its results for five minutes, that an ID field must be a UUID rather than a sequential integer, that calling a function with a deleted user's ID returns null. A useful comparison sits in something as small as a single line: a comment restating x = x + 1 as "increment x" adds nothing, because the code already says that. A comment that reads "offset for zero-indexed API response" earns its place, because no amount of staring at the line reveals that an external system's indexing convention is driving the arithmetic.

The same logic scales up to architectural decisions. Without that record, each new engineer who encounters the choice has to either trust it blindly or re-litigate it from scratch, and both outcomes cost time the team doesn't need to spend twice.

A familiar objection holds that self-documenting code, through well-chosen function and variable names, removes the need for most of this. That objection is partly right, and partly beside the point. AI tools fit into this picture more as a mechanical layer than a complete solution. What they cannot do reliably is supply the intent layer: the decisions that were made, the tradeoffs that were accepted, the constraints that shaped the design, and the alternatives that were considered and rejected. That layer still depends on an engineer who was present for the reasoning, or who took the trouble to ask.

The four distinct types of documentation

Much of what looks like a documentation quality problem is actually a category error: the wrong type of documentation applied at the wrong level, leaving readers to hunt through a reference page for a tutorial's hand-holding or through a tutorial for a reference page's precision. The Diátaxis framework, now widely adopted across technical writing, names four distinct orientations that documentation can take: tutorials, which are learning-oriented and walk a newcomer through a first successful experience; how-to guides, which are goal-oriented and solve a specific task for someone who already knows the basics; reference material, which is information-oriented and lays out facts without narrative; and explanation, which is understanding-oriented and gives context for why something works the way it does. Each orientation serves a different reader in a different mode, and conflating them produces documents that satisfy no one fully.

That same logic scales down into the practical layers a codebase actually contains. Docstrings capture function-level intent, written close enough to the code so it can travel with it. Inline comments handle line-level reasoning, the kind too granular for any higher document to carry. PR descriptions provide change-level context, explaining what a specific set of commits was trying to accomplish and why. Architecture Decision Records handle system-level decisions, the kind that outlive any single pull request. Each layer has a defined scope, and collapsing two of them into one artifact, a docstring trying to do an ADR's job, or a PR description trying to substitute for a reference page, produces noise.

One surface gets less credit than it deserves: configuration files written in YAML, JSON, or XML. A config file with no comments forces every future reader to either test each setting empirically or go find the person who wrote it.

Inline comments and the discipline of writing only what the code cannot say

Most inline comments fail because most of them restate what the code already shows. The discipline required is restraint paired with precision: restraint to avoid narrating every line, and precision to recognize the handful of lines where the code genuinely cannot express what a reader needs to know.

Self-documenting code, built from clear function names and expressive variables, reduces how many explanatory comments a function needs, but it does not eliminate the need for comments that capture constraints and decisions. The distinction that matters runs between narrating the code, which is unhelpful no matter how well-intentioned, and annotating what the code cannot show on its own, which is the entire point. TODO markers belong to this same inline layer, and they do real work when they name an owner or a specific condition under which the debt should be resolved. Left as permanent, unowned placeholders, they stop functioning as documentation and start functioning as clutter nobody feels responsible for clearing.

Architecture Decision Records as the place where the most valuable documentation lives

The biggest documentation gap on most engineering teams is missing engineering intent at the architectural level: the record of why a system is shaped the way it is. Architecture Decision Records exist to hold exactly that information, and a widely followed industry technology radar has recommended adopting them for years, keeping the practice in its top-recommended tier as recently as its most recent edition, a signal that this isn't a passing trend but an established discipline with years of institutional endorsement behind it.

The cost of skipping ADRs is visible in concrete incidents, not abstractions. In one documented case, a developer took almost two days to trace a critical billing function, purely because nobody had recorded why certain conditional branches existed. Missing context like that forces engineers to reconstruct history before they can safely act, turning what should be a quick fix into an investigation. An ADR addresses this by recording four things together: the decision itself, the context that made the decision necessary, the alternatives that were considered, and the reasoning for choosing one path over the others. That combination is what prevents a team from re-litigating the same architectural argument every time someone new joins or an audit requires justification for a system's design.

Missing engineering intent carries measurable downstream costs. Sprint timelines slip as missing context and undocumented assumptions force people to stop and ask around before they can proceed confidently.

Why do teams skip a practice with this much evidence behind it? The honest answer involves ownership, or the lack of it. The deeper issue is that documentation quality is rarely measured the way code quality is, so it has no structural defense when time runs short. A team can ship broken tests and get flagged immediately; it can ship an undocumented architectural decision and nothing stops the pipeline.

How Stripe made documentation quality a first-class engineering standard

Stripe's developer documentation is frequently cited as a model, and the reason holds up under scrutiny: Stripe treated its documentation with the same rigor most companies reserve for the product itself, rather than relegating it to a reference manual nobody budgets time for. A developer working in Python saw Python examples everywhere they went, with no manual translation required from a generic example into their own language's syntax.

What makes this case instructive is the institutional mechanism behind the polish of the output. Stripe built documentation quality into its engineering job ladders, making writing clear documentation a measurable, career-relevant skill. The company also ran internal classes on writing and documentation best practices, because it treated the skill as something you teach deliberately, not something you assume.

The business logic behind that investment is straightforward. Developers who can solve their own problems by reading documentation don't need a sales call to get started, don't need a scheduled onboarding session, and don't need to file a support ticket to understand basic integration behavior. Documentation built to that standard performs acquisition, onboarding, and support simultaneously, which converts what looks like a writing cost into a measurable reduction in a company's support burden. One might reasonably ask whether most teams could replicate the full scope of Stripe's investment. Probably not entirely in full scope. But the underlying mechanisms, documentation tied to job expectations and a definition of done that includes written context, scale down to teams of any size.

What sync between code and documentation requires

Documentation that was accurate on the day it was written but has since drifted from the code it describes is worse than no documentation at all, because it manufactures false confidence in a reader who has no way of knowing the gap exists. A developer who finds no documentation knows to proceed carefully. If a developer finds documentation that looks authoritative but describes a system that no longer exists, they often don't find out until something breaks.

If you treat documentation as something written once rather than maintained continuously, drift is the predictable result. It accumulates in silence, one small discrepancy at a time, until a developer acts on an instruction that used to be true and a regression follows. The failure mode runs parallel to undocumented edge cases causing production incidents: the documentation claims an edge case is handled a certain way, and the code has since moved on without telling anyone.

The docs-as-code approach treats this as a structural problem rather than a discipline problem, and addresses it accordingly. Documentation lives in version control alongside the code it describes, written in plain text formats that a diff tool can actually parse meaningfully. Automated checks catch broken links and style violations the way a linter catches code issues. CI/CD pipelines deploy documentation changes the same way they deploy code changes, keeping the two artifacts versioned together on the same schedule. The same definition-of-done mechanism that disciplined Stripe's engineering culture reappears here at the process level: a pull request cannot close without a corresponding documentation update, so the pipeline itself enforces synchronization.

How AI agents query documentation differently from humans

Documentation now has a second audience, and that audience reads differently than a human does. AI coding agents query documentation directly through retrieval systems rather than browsing it the way a developer scans a page for the relevant paragraph, and that shift changes what counts as good structure. Page scope, heading clarity, and self-containment stop being readability niceties and start determining retrieval accuracy: whether an agent pulls back the right answer or a plausible-sounding wrong one.

A page that tries to cover too much ground creates a specific failure for a retrieval system, even when a senior developer navigates that same page fluently. One documented case makes the failure concrete. The model began borrowing conventions from the wrong language entirely: code in one language started picking up parameter-passing patterns from another, an error a human reader would catch instantly but a retrieval system absorbed without noticing anything was wrong.

Short, self-contained pages with clear, specific headings give a retrieval model a cleaner target to match against a query. The unit of documentation needs to match the unit a retrieval system actually pulls back, rather than following whatever structure feels most natural to the person writing it. That reframes the argument against static, siloed documentation in structural terms rather than aesthetic ones: documentation not designed to be queried programmatically can't feed reliably into an agent's workflow, no matter how accurate its content is on the page. For any team building workflows around AI agents, documentation that stays synchronized with its codebase and is structured for retrieval is the mechanism by which an agent makes correct decisions about the system it's operating inside, and without that structure, accuracy at the sentence level doesn't translate into accuracy at the system level.

The properties that distinguish documentation a developer, or an agent, can act on

Across every layer this piece has worked through, from a single inline comment to a system-wide ADR to a page built for retrieval, a consistent set of properties separates documentation that functions from documentation that merely exists, and these properties can be tested.

Documentation that works captures intent rather than mechanism, recording the non-obvious constraint or the rejected alternative that the code has no way to express on its own. And it stays structured for retrieval, broken into self-contained, clearly headed units that a search system, human or automated, can actually match against a specific question.

None of these properties depend on volume. A team can write less and still produce documentation that holds up, provided what gets written captures the why that the code cannot say on its own. That is the standard this piece has been building toward from its first section: not more documentation, but documentation built to be acted on, by the next engineer who reads it and by the next system that queries it.

Sources

  1. Diátaxis
  2. Maintain an architecture decision record (ADR) - Microsoft Azure Well-Architected Framework
  3. Lecture Notes in Software Engineering Architectural Decision Records
  4. How Stripe creates the best documentation in the industry

More in Features