Most platform teams have an SLA for the services they run in production and nothing for the pipeline they run for their own engineers. The CI system is an internal product with users, a service level, and a failure mode that costs minutes per engineer per day. Treating its feedback latency as a favor — “CI is slow, we’ll look at it” — is why the complaint never closes. The fix is not a faster runner. It is a published number with an owner.
What an SLO actually is, and why CI qualifies
Google’s SRE book defines a service level indicator as “a carefully defined quantitative measure of some aspect of the level of service that is provided,” and notes that request latency — how long it takes to return a response — is commonly treated as a key SLI. An SLO is “a target value or range of values for a service level that is measured by an SLI,” naturally structured as SLI ≤ target or lower bound ≤ SLI ≤ upper bound. (Site Reliability Engineering, Chapter 4)
Nothing in that definition requires the service to be customer-facing. A CI pipeline is a request-driven system: an engineer pushes a commit, the system returns a pass/fail signal. The latency of that signal is a service level. If you already accept the SRE framing for your production API, you have no principled reason to exempt the pipeline.
The SRE workbook is blunt about the ownership requirement: “Once you have an SLO target below 100%, it needs to be owned by someone in the organization who is empowered to make tradeoffs between feature velocity and reliability.” (Implementing SLOs) For a CI SLO, that person is the platform team lead or the engineering manager who owns the developer experience budget. If no one can say no to a new blocking test stage, the number is decoration.
Why p95 and not the mean
The SRE book is direct about this: “averaging request latencies may seem attractive, but obscures an important detail: it’s entirely possible for most of the requests to be fast, but for a long tail of requests to be much, much slower. Most metrics are better thought of as distributions rather than averages.”
CI feedback time is exactly this shape. A monorepo with a warm cache and a small diff might return in three minutes. A full rebuild after a dependency bump might take forty. The mean sits somewhere in the middle and describes no one’s actual experience. The engineer who waited forty minutes did not experience the mean.
Percentiles let you describe the distribution’s shape. As the SRE book puts it, “a high-order percentile, such as the 99th or 99.9th, shows you a plausible worst-case value, while using the 50th percentile (also known as the median) emphasizes the typical case.” p95 is the pragmatic middle: it is the boundary of the bad day, not the catastrophe. It is the number you can defend in a planning meeting without arguing about outliers.
The workbook also supports multi-threshold latency SLOs — for example, “90% of requests are faster than 100 ms, and 99% of requests are faster than 400 ms” — which is worth considering if your pipeline has a bimodal distribution (cached vs. cold builds). But start with one number. p95 time-to-feedback is the one that maps to the engineer’s experience of “this build is taking forever.”
Define the SLI before you pick the target
The most common failure mode is picking a number before defining what is being measured. “Time-to-feedback” is not self-defining. You need to write down, in a document someone else can audit:
- Start event: the timestamp of the push or merge commit that triggered the pipeline. Not the time the runner picked up the job — that hides queue time, which is part of the engineer’s wait.
- End event: the timestamp of the first terminal signal the engineer can act on. If your pipeline posts a status check that says “failed” but the engineer has to open the logs to find out which stage failed, the end event is when the actionable signal appears, not when the last job exits.
- Denominator: every pipeline run triggered by a push to the default branch, including runs that were superseded by a later push. Excluding superseded runs is how you accidentally measure only the fast path.
- Numerator: runs whose end-to-start duration is at or below the threshold.
- Window: a rolling 28 days is the common SRE default; a calendar month is easier to report. Pick one and keep it.
This is the SLI specification. The implementation — which API you poll, which log field you parse — can change without changing the specification. The SRE workbook makes this distinction explicitly: an SLI specification is “the assessment of service outcome that you think matters to users, independent of how it is measured,” and the implementation is “the SLI specification and a way to measure it.”
Where the data comes from
You do not need a vendor feature called “CI SLO.” You need timestamps and a percentile calculation. The sources describe what is available:
GitHub Actions workflow run logs let you “see the time it took for each step to run,” and failed steps are automatically expanded. (GitHub Docs) The run-level start and end timestamps are available via the API; step-level timestamps let you attribute the tail to a specific stage.
GitLab’s Pipeline Insights provides “pipeline success and duration charts” with information about pipeline runtime and failed job counts. The same documentation notes that “external monitoring tools can poll the API and verify pipeline health or collect metrics for long term SLA analytics.” (GitLab Docs) That is the intended path: export the durations, compute the percentile yourself, keep the raw data.
Buildbot’s scheduler documentation describes a pattern that is directly relevant to feedback latency: “A quick scheduler might exist to give immediate feedback to developers, hoping to catch obvious problems in the code that can be detected quickly. These typically do not run the full test suite, nor do they run on a wide variety of platforms.” The same page documents a treeStableTimer that “will wait for this many seconds before starting the build” and restarts if new changes arrive. (Buildbot Docs) If your pipeline has a batching window, that window is part of time-to-feedback and must be in the SLI definition.
None of these tools ships a p95 time-to-feedback SLO dashboard. The claim that they do would be false. What they ship is the raw material: per-run and per-step durations, accessible via API or log export. The percentile is your job.
Set the target and the error budget
The SRE workbook warns against picking an SLO based on current performance because it “can commit you to unnecessarily strict SLOs.” But it also allows that current performance “can be a good place to start if you don’t have any other information, and if you have a good process for iterating in place.”
So: measure for two weeks. Compute p50, p95, and p99. Then set the SLO at a number that is achievable but not free. If your current p95 is 18 minutes, an SLO of 15 minutes is a real commitment. An SLO of 18 minutes is a description. An SLO of 10 minutes is a fantasy that will be ignored within a month.
The error budget is 100% minus the SLO. If the SLO is “95% of default-branch pipeline runs complete with an actionable signal within 15 minutes,” the error budget is 5% of runs. Over a 28-day window with 2,000 runs, that is 100 runs. When you burn the budget, something has to change: either the target moves, or the pipeline gets work. The SRE workbook is clear that the commitment to use the error budget for decision-making must be “formalized in an error budget policy.” Write it down. One page. Who is notified, what happens to the next sprint’s platform work, who decides.
Publish it where engineers already look
The SRE book makes the case for publishing: “Choosing and publishing SLOs to users sets expectations about how a service will perform. This strategy can reduce unfounded complaints to service owners about, for example, the service being slow.” It also notes the inverse: “Without an explicit SLO, users often develop their own beliefs about desired performance, which may be unrelated to the beliefs held by the people designing and operating the service.”
For a CI SLO, “publishing” means three things:
- A number in the README of the CI configuration repo. Not a dashboard link. The number, the window, and the current value.
- A weekly post in the engineering channel. One line: “p95 time-to-feedback this week: 14m 20s against a 15m SLO. Error budget: 62% remaining.”
- A breach trigger. When the budget is exhausted, the platform team opens a written post-mortem in the same format as a production incident: what burned the budget, what the contributing factors were, what changes are proposed, and what the new target is if the old one was wrong.
The post-mortem is the part most teams skip. It is also the part that makes the SLO real. A number that never triggers a written response is a number that will be ignored by the third month.
Keep it honest: pair it with a second metric
DORA’s guidance on software delivery metrics includes a warning that applies directly here: “Setting metrics as a goal. Ignoring Goodhart’s law and making broad statements like, ‘Every application must deploy multiple times per day by year’s end,’ increases the likelihood that teams will try to game the metrics.” (DORA Guides)
The gaming risk for a CI SLO is obvious: skip tests, move them to a non-blocking stage, or mark them as flaky and retry. The p95 number improves. The pipeline’s actual value to the engineer does not.
DORA’s recommendation is to “identify multiple metrics, including some with a healthy amount of tension between them.” For a CI SLO, the natural pair is a quality metric: change fail rate, or the rate of pipeline runs that pass but are followed by a revert or hotfix within 24 hours. If p95 time-to-feedback drops while the revert rate climbs, you have not improved the pipeline. You have moved the cost.
DORA’s own metrics — change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate — do not measure CI feedback latency. They measure delivery outcomes. A CI SLO is a leading indicator for those outcomes, not a substitute. Do not claim DORA compliance by publishing a p95.
The deprecation calendar applies to the pipeline too
GitLab’s pipeline efficiency documentation lists the factors that influence total pipeline duration: “Size of the repository, total number of stages and jobs, dependencies between jobs.” It also names a specific anti-pattern: “Tests that fail at the end of a long pipeline, but could be in an earlier stage, causing delayed feedback.” And it explains why ordering matters: “A job that takes a very long time to complete keeps a pipeline from returning a failed status until the job completes. Design pipelines so that jobs that can fail fast run earlier.”
These are not just optimization tips. They are the levers a platform team pulls when the error budget is burning. If a job no longer earns its place in the critical path — because the failure it catches is now caught by a faster check, or because it has not failed in six months — it should be moved to a non-blocking stage or removed. That decision should be visible in the same place as the SLO: a dated entry in the CI changelog with the reason and the measured effect on p95.
This is what it means to run the toolchain as an internal product. The pipeline has a service level, an owner, a budget, and a changelog. The deprecation calendar is not just for the tools you adopt. It is for the stages you keep.
A worked example, clearly labeled as hypothetical
Consider a 200-engineer monorepo where the p95 time-to-feedback is 22 minutes. The tail is driven by a single end-to-end test stage that runs last, after all unit and integration tests pass. The platform team defines the SLI as push-to-actionable-signal on the default branch, measures for two weeks, and sets an SLO of 15 minutes p95 with a 5% error budget.
They move lint and unit tests to an earlier stage that runs in parallel with the build, and split the end-to-end stage so that the fastest subset runs first. The p95 drops to 13 minutes. The error budget stops burning. The change is recorded in the CI changelog with the before-and-after numbers.
This example is constructed to illustrate the method. The numbers are not a recommendation and did not come from a real deployment. Your own baseline is the only valid starting point.
What to do on Monday
- Write the SLI specification in one paragraph. Start event, end event, denominator, window.
- Export the last 28 days of pipeline run durations. Compute p50, p95, p99.
- Pick a target that is 10–20% better than current p95. Write the error budget policy on one page.
- Name the owner. If you cannot name one person, stop here and fix that first.
- Publish the number in the CI repo README and post it weekly.
- Pair it with a quality metric so the number cannot be gamed by skipping tests.
- When the budget burns, write the post-mortem. When a stage stops earning its place, deprecate it in the changelog.
The pipeline is not a favor. It is a service. Publish the number.