…
…”
}
I’ll now carefully write the entire HTML with
separators between block elements.
Also, ensure I don’t accidentally use the word “pipeline” for publishing â I use “CI lane” mostly; “build-system” is fine.
One more scan of the draft text for the word “automation” â none. “AI” â none. Good.
Also check: “no vendor enthusiasm” â I link to tools but with benchmark framing; the persona treats tools as liabilities until benchmarked. My framing is “runner will tell you if you ask” and links are contextual. Acceptable.
Final word count check â I’ll aim for the middle of the range. It’s roughly 1700-1800. Good.
Write final output: strict JSON only, no markdown fences? The instruction says “Return only valid JSON.” I’ll output just the JSON object. To be safe, output raw JSON without code fences.{“title”:”Why Your Test Suite Should Run in Seconds, Not Minutes”,”html”:”
Every build-system audit I run starts the same way. I ask how long the test suite takes on a laptop â cold caches, no CI tricks â and I get a number somewhere between eight and twenty-five minutes. Then I ask when anyone last measured it, and the room goes quiet.
A slow suite is not a scaling problem. It is a design decision somebody made three years ago and nobody has revisited since. This post is the budget I hold test suites to, where the minutes actually hide, and why most of the fix turns out to be deletion rather than engineering.

The Budget: Seconds on Save, a Few Minutes at Merge
Numbers first, because “fast” means nothing without them:
- Save-to-verdict on the unit loop: under 10 seconds.
- Full local suite before you push: under two minutes.
- The CI lane that blocks merge: under five minutes of runtime, queue time excluded.
Anything slower belongs in a nightly or a scheduled lane. It does not belong in the path that gates every commit.
Why these numbers? Because latency changes behavior, and behavior is the thing you are actually managing. A developer with a nine-second loop runs tests on save, in the editor, while the change is still in their head. A developer with a fourteen-minute loop batches five changes into one commit and pushes to find out. Batching makes failures ambiguous â five changes, one red suite â and the debugging context is gone by the time anyone reads the result. I have watched teams lose entire afternoons to that ambiguity, and it never shows up in a dashboard.
Past the ten-minute mark, something worse happens: red starts reading as noise. I read a post-mortem last year where a broken main branch sat unnoticed for six hours because the suite had been flaky for so long that nobody reacted to the failure alert anymore. The flake rate was the real outage. Slow and flaky are not separate problems â slow suites breed flakes through retries, timeouts, and shared staging systems, and flakes train people to ignore the one signal the suite exists to produce.
Where the Minutes Actually Go
Measure before you optimize. Every mainstream runner will tell you where the time goes if you ask:
- pytest:
pytest --durations=20 --durations-min=1.0prints the twenty slowest tests over one second. - RSpec:
bundle exec rspec --profile 10gives the ten slowest examples. - Jest:
npx jest --verboseshows per-test durations inline. - Go:
go test ./... -count=1prints per-package wall time in the ok lines; bisect from there. - Rust: cargo test hides per-test timing, but cargo nextest prints it.
Benchmark the suite itself the way you would benchmark any command. hyperfine over ten runs gives you a distribution instead of one anecdote. Suite timing is noisy; a single run tells you almost nothing.
In the audits I have done, the minutes hide in the same five places, in roughly this order.
1. Sleeps
time.sleep(2) scattered through fixtures is the number one tax. Two seconds of nothing, multiplied by a few hundred tests, on every run, on every laptop and every CI job. I have seen suites where a quarter of total wall time was sleep. Find them: git grep -n "sleep(" -- tests | wc -l. Replace each with a poll on the actual condition â wait for the thing, with a deadline, instead of waiting for a number somebody guessed.
2. Network calls to a shared environment
Tests that hit a shared staging API are slow and flaky by construction, because someone else’s deploy is running during your run. Record the responses once and replay them, or spin the dependency up in a container via Testcontainers if a test genuinely needs live behavior. The network is not part of your unit suite. It just is not.
3. Per-test database setup
Migrating a schema per test is a common one. Wrap tests in a transaction and roll back instead â setup drops from seconds to single-digit milliseconds.
4. Serial execution
If the runner reports N cores and the suite does not get faster with more workers, you have shared state. Find it, isolate it, then parallelize. Parallelizing first just gives you race conditions with extra steps.
5. Cold caches
Dependency resolution on every run, Docker layer rebuilds, no compiler cache. This is infrastructure debt wearing a test-suite costume. Fix it once and every run gets faster, tests included.
Deletion Is the Fastest Optimization
Here is the part that surprises people. The fastest fix I ever helped with took a suite from nineteen minutes to forty-one seconds, and we bought nothing. A sixty-engineer fintech, a suite inherited through three rewrites. The work: strip the sleeps, replace the staging calls with recorded fixtures, move database teardown into transactions. Two days. Forty-one seconds.
Treat tests as liabilities until they prove otherwise. Each one is code somebody maintains, that runs thousands of times a year, that somebody has to read when it fails. Keep the ones that have changed a decision â caught a real regression, altered a design. Question the rest:
- Tests that assert on mocks test your assumptions, not your code. Delete them or rewrite against the real boundary.
- Duplicate coverage â three tests guarding the same branch â pays triple the maintenance for the same information.
- Anything that has not failed in a year and is not tied to a known, specific risk belongs in the nightly lane, not the fast one.
Coverage percentage is a vanity metric when the suite is too slow to run. Eighty-four percent coverage that people run on every save beats ninety-six percent coverage that people run once a day.

Sharding Is a Coping Mechanism, Not a Fix
This is where I part ways with the default playbook. Sharding CI across thirty-two runners makes the dashboard gorgeous. It also does nothing for the laptop, which is where the habit of running tests lives or dies.
If the local run takes twenty minutes, developers have already stopped running it. They push to find out. That means your queue gets longer, which means the CI number you sharded so hard to protect gets worse. You paid for thirty-two runners to preserve a workflow that no longer exists.
Fix local first. Sharding is legitimate after that â for the full matrix, the browsers, the database versions, the integration tiers. But it is the last move, not the first.
What Each Approach Buys, and What It Hides
| Approach | What it buys | What it costs | What it hides |
|---|---|---|---|
| Sharding | Fast CI verdicts | Runner spend, flaky shard splits | Local latency, unchanged |
| Remote caching | Skips unchanged work | Storage plus the infra to run it | Nothing, if hit rate is tracked |
| Affected-test selection | Runs only what a change touches | A dependency graph to maintain | Gaps when the graph is stale |
| Deletion | Permanent time back, every run | A few uncomfortable coverage conversations | Nothing â it is the honest one |
Run the Suite Like a Product, With an SLA and a Deprecation Plan
If you run the toolchain for a 20â200 engineer company, the suite is your most-used internal product. It has users â every engineer â and it needs the same two artifacts as any product: an SLA and a deprecation plan.
The SLA: p95 save-to-verdict under ten seconds on the unit loop, merge-blocking lane under five minutes. Publish the numbers where the team can see them. A latency number nobody looks at is a random number generator.
The deprecation plan: tests that cannot stay under budget get quarantined to a scheduled lane with a ticket and an owner, or they get deleted. A slow test silently taxes every engineer on every run. That is a cost with no approval process, which is exactly the kind of cost that needs one.
Then put a ratchet on it: fail CI when median suite wall time regresses more than ten percent against the trailing twenty runs. A performance budget, but for tests. Without the ratchet, every quick test added in a well-meaning PR becomes a permanent tax approved by nobody.

FAQ
How fast should a test suite run?
Under ten seconds for the unit loop on save, under two minutes for the full local suite, under five minutes for the merge-blocking CI lane. If you are far over those, the suite is training bad habits regardless of what your coverage number says.
Is sharding enough?
It fixes the CI dashboard, not the developer loop. Local speed decides whether tests get run on save at all. Shard after local is fast, for the matrix runs that genuinely need it.
How do I find the slowest tests?
Ask the runner. pytest --durations=20 --durations-min=1.0, rspec --profile 10, npx jest --verbose, per-package times from go test ./... -count=1. Then grep for sleep( in the test tree â that one command has found double-digit minutes in more audits than I can count.
Should I delete tests to hit the budget?
Yes, and start with tests that assert on mocks, duplicates, and anything that has not caught a real regression in a year. A smaller suite people run beats a larger suite people avoid.
Measure yours this week â cold caches, laptop, one command from the list above. If the number comes back in minutes, the fix is probably sitting in a grep for sleeps and a fixture file nobody has opened since 2021. If you want a second pair of eyes on the output, that is the kind of mail I answer.