Skip to content

B2B SaaS · CI on a hosted runner fleet · under NDA

A CI pipeline from 38 minutes to 9

A slow pipeline is rarely slow for the reason people assume. Ask a team where the time goes and the answer comes back confidently; measure it and the answer is usually somewhere else, and usually somewhere dull. So the first change was not a change at all.

The situation

A total duration tells you there is a problem and nothing about where it is. It also hides the most common cause: not the slowest step, but a fast step running far more often than it needs to. A step taking ninety seconds on every commit to every branch costs more per week than a step taking twelve minutes on merges to main, and it is invisible in any view that shows a single average.

Solution

01Measure the steps, not the pipeline

Per-step timing across a few hundred recent runs, kept long enough to see the variance rather than one sample. That separates three different problems that look identical in a total: steps that are reliably slow, steps that are usually fast and occasionally terrible, and steps that are fast but run everywhere.

02Make the caches actually hit

Nearly every slow pipeline already has caching configured. The question is the hit rate, and almost nobody measures it. A cache key built from something that changes on every commit is written every run and read never — pure overhead, and indistinguishable from a working cache in the configuration file.

The fix is to key on the thing that determines the contents, usually a lockfile hash, and then verify the hit rate afterwards rather than assume it. Same discipline as the rest of our work: the configuration states an intention and the measurement states the result, and only one of the two is evidence.

03Order the work so failure comes early

If a lint error can fail the build, it should fail it in the first thirty seconds rather than after the integration suite has run for twenty minutes. Reordering costs nothing and changes the felt speed of a pipeline more than most real optimisations, because the number that matters to a developer is time-to-answer, not time-to-green.

Then parallelism, but only where the dependency graph genuinely allows it, and sharding by historical duration rather than by file count so the shards finish together.

Outcome

38
minutes before
9
minutes after
0
checks removed

Every check that was there before is still there — nothing was removed to make the number look better, and the CI provider did not change. What changed was how often the expensive work runs, whether the caches were doing anything, and the order of the cheap checks.

Core tech

  • GitHub Actions
  • GitLab CI
  • Docker
  • Argo CD
  • Terraform

Method

What makes this routine rather than risky

Nothing here is proprietary. This is the part that transfers to your estate whether or not you ever call us.

Per-step timing before any change

Across enough runs to see the variance, not one sample. A step that is usually fast and occasionally terrible is a different problem from one that is reliably slow.

Verify the cache hit rate

A cache keyed on something that changes every commit is written every run and never read. It looks correct in the config and does nothing.

Run expensive work only where it pays

Full suites on merges, fast feedback on branches. Most pipelines run everything everywhere because that was the safe default on day one and nobody revisited it.

Fail fast, deliberately

Cheap checks first. The metric that matters is how long the developer waits for an answer, not the pipeline's total duration.

Shard by duration, not by file count

Balanced shards finish together. Shards split by file count are only as fast as their unluckiest member.

Change the pipeline, not the vendor

Migrating CI providers is a project with its own risk, and it does not fix a cache key. The provider is rarely the actual constraint.

The method on this page is standard practice and checkable. The specific figures are illustrative pending the engagement record.