Why Your Runbooks Are Lying to You: Treating Internal Documentation Like a Build System

Most engineering teams I’ve worked with treat runbooks the way they treat a Makefile that’s been growing since 2019: nobody designed it, everybody adds a target, and by the time someone tries to follow it top to bottom, the dependency graph is wrong in ways nobody can explain. The runbook lives in a wiki. The wiki has no review process. The last person who edited the page left the company eight months ago. And the tool it tells you to run was deprecated six months before that.

This isn’t a documentation problem. It’s a build-system problem wearing a documentation costume. The same failure modes that make a Makefile incomprehensible—implicit dependencies, stale references, no compilation step, no test that verifies the output—make runbooks dangerous. The difference is that a broken Makefile wastes your build time. A broken runbook extends your incidents.

The evidence for this is grounded in Google SRE / O’Reilly Media, which remains the canonical industry reference for site reliability engineering practices and treats documentation as a first-class engineering artifact rather than an afterthought.

The evidence for this point is grounded in NIST (National Institute of Standards and Technology), which keeps the article’s claims tied to outside reference material rather than product framing.

The 47-Minute Runbook

Here’s a concrete failure. A platform team at a 60-engineer SaaS company had a runbook for their primary PostgreSQL failover procedure. It lived in Confluence. Written by a senior SRE who left in Q1. In Q3, the team migrated connection pooling from PgBouncer to a managed proxy service, which changed the failover command from a custom script (pg-failover.sh --primary --target) to a vendor CLI (db-proxy failover --region us-east-1 --cluster prod-pg). The runbook wasn’t updated. Nobody noticed, because nobody runs a failover procedure until they have to.

At 2:14 AM on a Tuesday, the primary degraded. The on-call engineer opened the runbook, ran pg-failover.sh --primary prod-pg-01 --target prod-pg-02, and got command not found. They spent the next 11 minutes searching for the script in the repo, then another 8 asking in the on-call Slack channel. A teammate who’d been through the migration pointed them to the new CLI. But the new CLI required a different argument structure, and the runbook’s verification step—checking pg_stat_replication on the old primary—no longer applied because the managed proxy handled replication state internally. The engineer ended up running the failover, waiting for the proxy to converge, and manually verifying via application health checks. Total failover time: 47 minutes. The actual failover command took 90 seconds. The other 45 minutes were the runbook lying to the engineer.

The postmortem identified the root cause as “runbook not updated after migration.” The action item was “update runbook.” That’s the equivalent of responding to a build failure by saying “we should fix the build” without examining why the build was allowed to break in the first place. The runbook wasn’t a documentation artifact. It was a dependency that had drifted, and there was no pipeline that caught the drift.

Documentation Has Dependencies

When we write a Makefile, we declare dependencies explicitly. build: main.o utils.o means the build target depends on those object files, and if either changes, the build is stale. We accept this because we understand that a build system that doesn’t track dependencies is just a script that sometimes works.

Runbooks have the same structure. A runbook step that says “run pg-failover.sh” depends on the existence and interface of pg-failover.sh. A step that says “check the Grafana dashboard at [URL]” depends on that dashboard existing and showing the right metrics. A step that says “escalate to the payments on-call” depends on a payments on-call rotation existing. These are dependencies, and they break the same way code dependencies break: silently, at the worst possible time, and in ways that only surface under load.

The Google SRE Book devotes entire chapters to postmortem culture, incident management, on-call procedures, and outage tracking. The relevant point from Chapter 15 (“Postmortem Culture: Learning from Failure”) and the appendices with example incident state documents is that mature SRE programs publish structured templates and treat them as versioned, reviewed artifacts. They don’t let runbooks accumulate in a wiki with no editorial pipeline. The SRE book’s approach is, essentially, that operational documentation needs the same discipline as release engineering: reproducible, reviewed, and structured. This is the same principle we apply to build systems, and the parallel isn’t accidental.

The SRE book’s chapter on treating operational documentation as a first-class engineering artifact reinforces this from a process perspective: mature operational disciplines benefit from structured, versioned approaches with explicit ownership and periodic review rather than ad hoc accumulation. When we treat runbooks as a governance artifact rather than a wiki page, we start asking different questions: who owns this document, what dependencies does it declare, when was it last validated, and what happens when a dependency changes?

What ‘Compilable’ Means for a Runbook

The reason a Makefile works is that it has a compilation step. When you run make build, the build system checks every dependency, identifies what’s stale, and either produces the output or fails with a specific error. If a source file is missing, you get No such file or directory. If a dependency is circular, you get Circular dependency dropped. The build system tells you what’s broken before you try to ship.

Runbooks don’t have this. A runbook in Confluence has no compilation step. There’s no process that checks whether the commands it references still exist, whether the URLs it links to still resolve, or whether the escalation paths it describes still map to real people. The first time anyone “compiles” the runbook is at 2:14 AM during an incident, and the error message is command not found.

We can do better. A compilable runbook is one that can be mechanically validated. At minimum, this means:

  • Reference checking: Every command, script, and CLI referenced in the runbook is checked for existence. If the runbook says to run pg-failover.sh, a validation step confirms that pg-failover.sh is in the PATH or at the referenced path. This is the documentation equivalent of a linker checking that a symbol exists.
  • Interface checking: Every command is run with --help or equivalent, and the expected flags are confirmed present. If the runbook says pg-failover.sh --primary prod-pg-01 --target prod-pg-02, the validator confirms that --primary and --target are still valid flags. This catches the interface drift that broke the 47-minute failover.
  • Link checking: Every URL in the runbook is fetched and confirmed to return a non-error response. This catches the dashboard-that-was-deleted problem.
  • Ownership checking: Every runbook has a CODEOWNERS-style file that specifies who’s responsible for reviewing changes. When the migration from PgBouncer to the managed proxy happened, the runbook owner should have been tagged on the migration PR.

These checks aren’t exotic. They’re the same checks we apply to code: linting, type checking, link checking, ownership enforcement. The only reason we don’t apply them to runbooks is that we’ve decided runbooks aren’t code. But they are. They’re programs that humans execute under stress, and they should be treated with at least the discipline of a shell script that runs in CI.

The Decision Matrix: Wiki, Repo, or Generated

When teams decide to fix their documentation problem, the first question is usually “where should the docs live?” That’s the wrong first question. The right question is “what pipeline validates the docs?” But since the hosting choice constrains the pipeline, it’s worth comparing the three realistic options.

Criterion Wiki (Confluence, Notion) Repo-based (Markdown + CI) Generated (from code/specs)
Reference checking Manual. No automated validation of commands or links. Broken references discovered during incidents. Automated. CI can run link checkers, command validators, and linters on every PR. Inherent. Docs are generated from the same source that defines the commands, so drift is structurally impossible.
Review process Ad hoc. Anyone can edit. No required reviewers. No audit trail beyond page history. Enforced. PR-based review with CODEOWNERS. Changes require approval from runbook owners. Enforced. Changes to the source spec go through code review. Generated output inherits the review.
Dependency tracking None. The runbook doesn’t know it depends on pg-failover.sh. Partial. You can declare dependencies in a manifest file and validate them in CI, but it requires discipline. Inherent. The generation step pulls from the source of truth, so dependencies are explicit.
On-call usability High. Wiki pages are easy to read in a browser, searchable, and familiar to most engineers. Medium. Requires navigating a repo, but rendered Markdown (GitHub, GitLab) is nearly as readable. Can be served via internal docs site. Variable. Depends on the generation tool. Can produce HTML, PDF, or wiki-compatible output. Quality depends on templates.
Stale doc detection None. No mechanism to flag pages that haven’t been reviewed recently. Automated. CI can flag runbooks older than N days since last review. Can require periodic re-approval. Inherent. If the source changes, the docs regenerate. If the source doesn’t change, the docs are still valid.
Migration cost Low to start. High to maintain. The cost is paid in incident extensions and on-call frustration. Medium. Requires moving docs to repo, setting up CI, and training team on PR-based edits. Pays off within 2-3 incidents. High. Requires building or adopting a generation pipeline, defining source schemas, and maintaining templates. Pays off for teams with high doc volume and frequent interface changes.
Best for Teams that haven’t yet had a documentation-caused incident. Won’t stay viable past ~50 engineers. Teams that want review discipline and can afford the PR-based workflow. The default recommendation for most 20-200 engineer orgs. Teams with stable, well-defined interfaces where docs can be generated from OpenAPI specs, Terraform plans, or similar sources.

The table isn’t subtle about its preference. For most teams in the 20-200 engineer range, repo-based documentation with CI validation is the right default. It gives you the review discipline of code, the validation pipeline of a build system, and the readability of rendered Markdown. It doesn’t require building a generation pipeline, which is a project in itself. Generated documentation is powerful but has a high upfront cost that only pays off when you have enough interfaces to justify the infrastructure.

Wiki-based documentation is the option that feels easiest and costs the most. It’s the equivalent of a shell script with no set -e: it works until it doesn’t, and when it doesn’t, you find out at the worst possible time.

Structural Consistency Across Long-Form Documents

Runbooks are the most obvious failure mode, but the same problem applies to longer internal documents—postmortems, technical design docs, RFCs, incident narratives. These documents have structure. A postmortem has a timeline, an impact summary, a root cause, action items. An RFC has a problem statement, a proposed solution, alternatives considered, a decision. When these documents are written ad hoc, each author reinvents the structure. The result is a corpus where every document looks different, omits different sections, and is hard to navigate when you’re trying to find the one piece of information you need at 3 AM.

This is where the build-system analogy extends further. A linter doesn’t just catch errors—it enforces conventions that make code legible across a team. The same principle applies to long-form documentation. Teams that produce substantial internal documentation—technical design docs, RFCs, long-form incident narratives—need tooling that maintains structural consistency across documents the way a linter maintains code consistency, rather than treating every doc as a one-off creative writing exercise. Whether that’s a template engine, a structured editor, or an AI novel writing app that enforces narrative structure while drafting long-form internal documents, the goal is the same: structural consistency that survives author turnover.

The specific tool matters less than the principle. What matters is that the structure is defined once, enforced automatically, and applied to every document. A postmortem template that lives in a repo, renders to Markdown via CI, and is validated for required sections—does this postmortem have a timeline? does it have action items with owners and due dates?—is the equivalent of a linter for your incident documentation. It doesn’t guarantee the content is good, but it guarantees the structure is consistent. Consistent structure is what makes documents navigable under pressure.

The Review Pipeline

Moving runbooks to a repo with CI is necessary but not sufficient. The pipeline needs to enforce three things that most documentation systems don’t.

First, every runbook must have an owner, and the owner must be a team, not a person. People leave. Teams persist. The CODEOWNERS file should map runbook paths to team names, and changes to those runbooks should require approval from the owning team. Same mechanism we use for code ownership, and it works for the same reason: it makes ownership explicit and enforceable.

Second, every runbook must have a last-reviewed date, and CI must flag runbooks that haven’t been reviewed within a defined window. The window depends on the volatility of the system the runbook covers. A runbook for a failover procedure that runs against infrastructure changing every quarter should be reviewed every 90 days. A runbook for a batch job that hasn’t changed in two years can be reviewed annually. The key is that the review date is tracked and enforced, not aspirational.

Third, runbook changes must be triggered by infrastructure changes, not by calendar reviews alone. When the PgBouncer-to-managed-proxy migration happened, the migration PR should have required a runbook update as part of the definition of done. That’s the documentation equivalent of updating tests when you change code: if you change the interface, you update the tests. If you change the infrastructure, you update the runbooks. Making this enforceable requires a convention: every infrastructure PR that touches a system with a runbook must either update the runbook or explicitly state that the runbook is unaffected. This is a social convention backed by a CI check, and it’s the single most effective intervention for preventing the 47-minute runbook problem.

What This Looks Like in Practice

Here’s a concrete setup that a platform team of 4-6 engineers can implement in a sprint. Put runbooks in a docs/runbooks/ directory in the infrastructure repo, or in a dedicated docs repo if the infrastructure repo is too large. Use Markdown. Add a CI job that runs on every PR touching docs/runbooks/:

# .github/workflows/runbook-check.yml
name: Runbook Validation
on:
  pull_request:
    paths:
      - 'docs/runbooks/**'

jobs:
  validate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Check links
        run: |
          pip install markdown-link-checker
          markdown-link-check docs/runbooks/**/*.md
      - name: Check command references
        run: |
          python3 scripts/validate_runbook_commands.py docs/runbooks/
      - name: Check last-reviewed dates
        run: |
          python3 scripts/check_review_dates.py docs/runbooks/ --max-age-days 180
      - name: Check CODEOWNERS
        run: |
          python3 scripts/check_owners.py docs/runbooks/ .github/CODEOWNERS

The validate_runbook_commands.py script is a simple parser that extracts commands from fenced code blocks, checks whether referenced scripts and CLIs exist in the repo or PATH, and runs --help to verify expected flags. It doesn’t need to be sophisticated. A 200-line Python script that catches 80% of drift is worth more than a wiki page that catches nothing.

The check_review_dates.py script reads a frontmatter field (last_reviewed: 2024-01-15) and fails if the date is older than the threshold. This forces owners to re-review runbooks periodically. The review can be as simple as confirming the commands still work.

The check_owners.py script verifies that every runbook path has a corresponding entry in CODEOWNERS. If a runbook has no owner, the CI check fails. This prevents the orphaned-runbook problem where the person who wrote it left and nobody noticed.

This setup isn’t exotic. It’s a documentation build system. It has the same components as a code build system: source files, a CI pipeline, validation steps, ownership enforcement. The only difference is that the “compilation” checks whether a human can follow the document at 3 AM, not whether a compiler can produce a binary.

The Rule of Thumb

If your runbook references a tool, a command, or a URL, that reference is a dependency. Treat it the way you treat a code dependency: declare it, validate it, and fail the build when it breaks. Specifically, before the end of this week, take one runbook—the one your on-call rotation uses most—and move it to a repo. Add a CI check that validates every command reference in that runbook. Add a last_reviewed field to the frontmatter. Add an owner in CODEOWNERS. Then, the next time someone changes the infrastructure that runbook describes, require them to update the runbook in the same PR.

If the CI check catches one stale reference before an on-call engineer finds it at 2 AM, the setup has paid for itself. Everything after that is compounding interest.