CI Performance — Caching, Parallelism & Flaky Jobs
CI caching layers, parallelism/sharding, flake economics, and metrics that actually drive feedback time.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Pull-request CI sits at 25 minutes cold and 8 minutes warm
Prefer
Make the warm path correct
Key caches so a hit is the common case, and a hit cannot go green for the wrong bytes. Measure hit rate.
- Lockfile plus runtime version plus toolchain is the key.
- A wrong hit is worse than a cold install.
- Bigger runners come after the compile is actually CPU-bound.
Alternative
Buy larger runners first
The wall clock drops until the cache misses, the matrix fans out, or a dirty workspace poisons the next job.
- Spend tracks idle matrix cells.
- Sticky self-hosted disks leak state across jobs.
- You still do not know the hit rate.
From pull request to a signal you can trust
The source diagram branches on lockfile, test strategy, and flake. This column is the same path: key, restore or install, compile, shard or select, then quarantine or green.
- 1
Key the cache
Lockfile hash, OS, language major, and tool version. A lockfile change is a cold install, then a save. - 2
Compile from inputs
Remote or local build cache hits only when the action inputs match. Non-hermetic tools poison hits. - 3
Spend parallelism on the slow part
Shard by historical timing, or run the affected graph. Merge the reports. - 4
Treat flakes as a budget
Cap retries. Track the rate. Hand test isolation to the testing cluster with an owner.
Overview
Slow CI is a product tax. People batch work, skip local checks, and merge larger diffs. Flaky CI is worse: the team clicks re-run until the badge is green.
Performance work is cache correctness, parallelism with isolation, and flake economics. It starts as job design. "Buy a bigger runner" is a later sentence, after queue wait and run time say the machine is the bottleneck.
p95 pull-request feedback past 20 to 30 minutes costs a context switch. A wrong cache (stale compiler output, a node_modules tree from another Node major) causes bugs that disappear on a clean runner. Parallelism without shard isolation multiplies races. Those races are Flaky Tests. Here the question is how the job is cut.
Caching layers
| Cache | Speeds up | Invalidation key | Footgun |
|---|---|---|---|
Dependency (npm, pip, gradle) | Install | Lockfile hash | node_modules reused across Node majors |
Build (Bazel, Gradle remote, ccache) | Compile | Action inputs | Non-hermetic tools return wrong hits |
| Docker layer cache | Image build | Dockerfile instruction order | COPY . . early busts every later layer |
| Remote registry cache | Pull and push | Digest | Unauthenticated public caches |
| Test result / affected tests | Test runtime | Content-hash graph | A wrong graph under-tests the change |
A warm 8 minute job is a success only when warm is the common case and the hits are correct. Track hit rate, and track incidents where a cache change produced a wrong green.
Caching node_modules across Node majors is a classic miss: native addons and the resolved tree differ, and the failure looks like an application bug. Key by runtime version plus lockfile.
Flow
- 1
1. Pull request arrives
- next2. Key the cache on lockfile
- 2
2. Key the cache on lockfile
- next3. Restore or install cold
- 3
3. Restore or install cold
- next4. Reuse the compile cache
- 4
4. Reuse the compile cache
- next5. Shard or select tests
- 5
5. Shard or select tests
- next6. Merge the reports
- 6
6. Merge the reports
- next7. Quarantine or go green
- 7
7. Quarantine or go green
Lesson map
CI Performance — Caching, Parallelism & Flaky Jobs
CI caching layers, parallelism/sharding, flake economics, and metrics that actually drive feedback time.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Pull request arrives"] b["2. Key the cache on lockfile"] c["3. Restore or install cold"] d["4. Reuse the compile cache"] a -->|1. Pull request arrives| b b -->|2. Key the cache on lockfile| c c -->|3. Restore or install cold| d
Parallelism
- Matrix. OS times language version. Useful for libraries. Cap the fan-out. Kill axes that never fail.
- Test sharding. Split by historical duration, not by file count. A shard of tiny files and a shard that holds the 40 second suite is not parallelism.
- Path filters. Docs-only pull requests skip the image build. The cost is a missed build when generated code or docs affect runtime. Prefer an explicit dependency graph over a naive ignore list. Under-approximation misses regressions. Over-approximation is slower and safer. Start conservative.
- Merge queue. Spend machines to keep
maingreen. The anatomy page owns where that sits in the graph. - Remote execution. Hermetic builds (Bazel and similar) scale horizontally because inputs are declared. That declaration is the hermetic-build note on the artifacts page.
Aggressive sharding cuts wall-clock and raises utilization. Each shard pays setup cost. Global fixtures contend. Racy tests surface more often. The bill goes up. Parallelize a database suite with a schema or database per shard, or with transactional isolation. One shared database and racing TRUNCATE is a flake factory. The isolation technique is the testing cluster. The job boundary is this page.
Flaky jobs
A flaky job is a flaky test or flaky infrastructure: a spot runner, a registry blip, a package-host 503.
| Source | What to change in CI |
|---|---|
| Tests | Quarantine with an owner. Isolation, clocks, and order live in Flaky Tests |
| Package registries | Retry with backoff, plus an internal mirror |
| Disk full on a shared runner | Ephemeral cleanup, larger disks. Do not depend on a dirty workspace |
| Cache corruption | Version the key. Verify a checksum before you trust the hit |
| Timing in deploy smoke | Health checks with a deadline. A fixed sleep hides the race |
Cap retries. Track flake rate. Page an owner. "Retry until green" with no cap is how a 1 percent flake becomes the team's afternoon.
A workable SLO is under 1 percent job retries on the default main pipeline. Hotter than that, the suite is quarantined and owned. Two hundred pull requests a day, ten jobs, 1 percent flakes: about twenty noisy failures a day. Five minutes of triage each is more than an hour and a half of engineering, daily, on nothing. Quarantine pays for itself.
What to measure
| Metric | Why it matters |
|---|---|
| p50 / p95 time to first failure | Interrupt cost while the author is still in the change |
| p50 / p95 time to green | Merge latency |
| Cache hit rate by key | Whether the warm path is real |
| Flake rate (retries) | Whether red still means broken |
| Cost per merged pull request | Spend sanity, including idle matrix cells |
| Queue wait versus run time | Runner capacity versus job design |
If queue wait dominates, add runners. If run time dominates, fix keys, shards, and tests.
Runner topology
| Topology | Pros | Cons |
|---|---|---|
| GitHub-hosted | No runner ops, clean VMs | Cold caches, minute caps |
| Self-hosted warm | Cache locality | Patching, sticky state, a larger security surface |
| Remote build cluster | Compile scales out | You pay the hermeticity work up front |
Warm self-hosted runners that reuse a workspace can leak secrets across jobs. Prefer ephemeral VMs or containers even when they cost a few seconds. Bigger runners beat better caching when the compile is truly CPU-bound and the cache hit rate is already high. Otherwise fix keys and hermeticity first.
Docker build tips
- Order the Dockerfile for layer reuse. Dependencies land before the application copy.
- Use BuildKit cache mounts for package managers.
- Copy lockfiles before the full build context.
- Multi-stage builds: the digest you promote is the runtime stage.
- An
apt-get updatewithout pinned versions fights reproducibility. - Push and pull from a registry in the same region as the runners.
Dependency cache and a path filter
on:
pull_request:
paths-ignore:
- "**/*.md"
- "docs/**"
jobs:
unit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- run: pip install -r requirements.txt
- run: pytest -qpytest -n auto (xdist) belongs here only after each worker has its own database or schema. The path ignore is a sketch. A generated client that lives under docs/ will skip the image build. The helper below is the stricter check.
Helpers you can run
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Why is caching node_modules across Node majors dangerous?
Answer
Native addons and the resolved tree differ. The crash looks like product code. Key the cache by runtime version and lockfile.
Cold CI is 25 minutes and warm CI is 8. Is that a success?
Answer
Only if warm is the common case and the hits are correct. Measure hit rate. Track wrong-green incidents after a cache change.
How do you parallelize a Django or Postgres suite safely?
Answer
Give each shard its own database or schema, or isolate with transactions. A shared database plus racing truncates will flake. How to freeze time and quarantine the test is Flaky Tests.
What is the cost of path filters?
Answer
A codegen or docs path that affects runtime never builds. Prefer an explicit dependency graph. Missing a regression is worse than a slow extra job.
When do bigger runners beat better caching?
Answer
When the compile is CPU-bound, hermetic, and the cache hit rate is already high. Otherwise fix keys first.
How do you budget CI spend?
Answer
Cost per merged pull request, cache storage, and idle matrix cells. Delete low-signal matrix axes. If queue wait dominates the trace, add runners. If run time dominates, fix the job.
What flake-rate SLO would you defend?
Answer
Under 1 percent retries on the default main pipeline is a concrete example. Anything hotter gets a quarantine ticket with an owner. Uncapped re-runs hide the rate.
Pitfalls
- A cache key of
mainthat every branch shares, including broken dependency trees. - Sharding by file count so one shard holds every slow test.
- Path filters copied from a blog, then a generated SDK never rebuilds.
- Self-hosted runners that keep
~/.npmand yesterday's secrets. - A dashboard of total CI minutes with no split between queue wait and run time.
Pick a 20 minute pull-request job. Write the cache key, the shard count, and the one path that must still build the image. Then write the flake budget for a week if 1 percent of jobs retry and each retry costs five minutes.