Build Systems & Monorepos - Graphs, Caching, Hermeticity & Reproducible Environments
Concept hub: a build is a function over a dependency graph; the four properties (correctness, incrementality, caching, hermeticity) and how they stack; task runner vs build system vs package/environment manager; what Turborepo, Nx, Bazel and Nix share and where they differ; runnable mini build function and graph-granularity demo; decision chart.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a build system compute?
Answer
Outputs from declared inputs, over a dependency graph, doing as little work as possible.
L2
Name the four properties on this page.
Answer
Correctness, incrementality, caching and hermeticity.
L3
Task runner vs build system?
Answer
A task runner orders and caches coarse package scripts. A build system owns fine-grained actions with declared inputs and outputs.
L4
Why do caching and hermeticity go together?
Answer
A cache key hashes declared inputs. If a task reads undeclared state, the same key can map to different outputs.
L5
What do Turborepo, Nx, Bazel and Nix share?
Answer
Hash inputs, look up a cache, skip identical work. They differ in the unit of work and how precisely inputs are known.
L6
When is Turborepo enough and when do you reach for Bazel?
Answer
Turborepo fits JS/TS package scripts. Bazel or Buck2 fit large polyglot repos that need remote execution and cannot tolerate wrong cache hits.
L7
Why is Docker a weak incremental build graph?
Answer
Layers form a linear chain, so any change invalidates every later layer.
Failure modes
An undeclared env var produces a wrong cache hit
The key looked the same, but a value the task read was different, so the cache served an output built from other inputs.
Undeclared outputs vanish on a cache hit
A task runner only restores files it was told about, so a cached build can succeed with missing files.
Volatile stamping defeats caching
Embedding the time or git SHA into every build changes every downstream key.
Misconceptions
A cache hit proves the build is correct.
A hit only means the key matched. If the key misses a real input, the hit is wrong.
Turborepo, Nx, Bazel and Nix are interchangeable.
They share the hash-and-skip idea but differ in unit of work, how the graph is known, and hermeticity.
File modification times are a reliable change signal.
touch, checkouts, clock skew and backup restores change mtimes without changing content, or the reverse.
Interviewer traps
Recommending Bazel for a ten-package TypeScript repo.
Start with a task runner and migrate when correctness or scale demands it.
Calling Docker your build system.
Docker is good for packaging and environments, weak as an incremental build graph.
Design scenario
Same prompt for every reader.
Requirements
Faster CI on most PRs, no wrong cache hits in production builds, and a migration the team can finish in weeks.
Failure assumptions
- Some scripts read env vars that are not declared.
- A few tasks write outputs that are not listed.
- Root config files change several times a week.
Constraints
- Do not rewrite every build as Bazel rules up front.
- Production builds must not reuse preview artifacts.
Prompt
A 40-package TypeScript monorepo has 25-minute CI and occasional 'works after rm -rf dist' failures. Pick the build approach and the first three changes.
API
Which task inputs and outputs do you declare first?
Data
What goes into the cache key for the production bundle?
Architecture
Where do the task runner, remote cache and lockfile sit, and when would you revisit Bazel?
Overview
Turborepo, Nx, Bazel, Make, Gradle and Nix look like very different tools, but they all answer the same question: given these inputs, what outputs should exist, and what is the least work needed to get them? A build system is a function from inputs (source files, toolchains, flags, environment, dependencies) to outputs (binaries, bundles, test results), evaluated over a dependency graph. Each tool picks a different point on four axes: correctness (outputs always match a clean rebuild), incrementality (redo only what changed), caching (reuse results by a key, even across machines) and hermeticity (a task can only see what it declared). Learn those four ideas and every tool on this page becomes a set of trade-offs instead of a new thing to memorize.
A build is a function over a graph
The paper Build Systems a la Carte (Mokhov, Mitchell and Peyton Jones, ICFP 2018) gives the cleanest vocabulary:
- Store: a map from keys to values. In a software build the keys are file paths and the values are file contents.
- Task: how to compute one key from the values of other keys (compile
main.cintomain.o). - Build system: takes the tasks, a target and the current store, and returns a store where the target and everything it depends on is up to date.
- Correct means: inputs were not corrupted, and recomputing any output from the final store gives exactly the value already there.
- Minimal means: each task runs at most once per build, and only if something it transitively depends on changed.
Everything else (remote caches, sandboxes, affected detection, /nix/store) is an engineering answer to "how do we stay correct while doing as little work as possible, for as many people as possible?"
The four properties
| Property | Question it answers | How tools get it | What breaks without it |
|---|---|---|---|
| Correctness | Does the output equal what a clean build would produce? | Complete dependency info, content hashing, sandboxing | "Works after rm -rf dist", stale binaries, flaky CI |
| Incrementality | Can we redo only what changed? | Dependency graph plus change detection (mtime or hash) | Every edit triggers a full rebuild |
| Caching | Can we reuse a result computed earlier, or by someone else? | Cache key = hash of all declared inputs; local or remote store | Teams and CI repeat identical work |
| Hermeticity | Can a task see anything it did not declare? | Sandboxes, pinned toolchains, filtered env vars | Cache keys miss real inputs, so caches return wrong results |
The properties stack. Incrementality without correct dependencies gives stale output. Caching without hermeticity gives wrong cache hits: the key looked the same, but an undeclared input (an env var, /usr/bin/gcc, the time of day) was different. That is why the most aggressive cachers (Bazel, Buck2, Nix) are also the strictest about declared inputs.
Decisions
- 1
Step 1: read the task graph (targets, inputs, commands)
- nextStep 2: topologically order tasks and find independent ones
- 2
Step 2: topologically order tasks and find independent ones
- nextStep 3: hash each task's declared inputs into a cache key
- 3
Step 3: hash each task's declared inputs into a cache key
- nextStep 4: key found in local or remote cache?
- ?
Step 4: key found in local or remote cache?
- nextStep 5a: restore outputs and logs, skip the work
- nextStep 5b: run the task, ideally in a sandbox
- 5
Step 5a: restore outputs and logs, skip the work
- nextStep 7: store outputs under the key and continue downstream
- 6
Step 5b: run the task, ideally in a sandbox
- nextStep 6: did the task read an undeclared input?
- ?
Step 6: did the task read an undeclared input?
- nextStep 7: store outputs under the key and continue downstream
- nextFailure path: cache entry depends on hidden state, later hits can be wrong
- 8
Step 7: store outputs under the key and continue downstream
- 9
Failure path: cache entry depends on hidden state, later hits can be wrong
- nextFix: declare the input, sandbox the task, or mark it uncacheable
- 10
Fix: declare the input, sandbox the task, or mark it uncacheable
Lesson map
Build Systems & Monorepos - Graphs, Caching, Hermeticity & Reproducible Environments
Concept hub: a build is a function over a dependency graph; the four properties (correctness, incrementality, caching, hermeticity) and how they stack; task runner vs build system vs package/environment manager; what Turborepo, Nx, Bazel and Nix share and where they differ; runnable mini build function and graph-granularity demo; decision chart.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Step 1: read the task graph (targets, inputs, commands)"] b["Step 2: topologically order tasks and find independent ones"] c["Step 3: hash each task's declared inputs into a cache key"] d["Step 4: key found in local or remote cache?"] e["Step 5a: restore outputs and logs, skip the work"] f["Step 5b: run the task, ideally in a sandbox"] g["Step 6: did the task read an undeclared input?"] h["Step 7: store outputs under the key and continue downstream"] x["Failure path: cache entry depends on hidden state, later hits can be wrong"] y["Fix: declare the input, sandbox the task, or mark it uncacheable"] a -->|continues| b b -->|continues| c c -->|continues| d d -->|continues| e d -->|continues| f f -->|continues| g g -->|continues| h g -->|continues| x e -->|continues| h x -->|continues| y
How precisely does the tool know its inputs?
Prefer
Declare every input, then cache aggressively
The more precisely a tool knows each task's inputs, the more safely it can skip, cache and distribute work.
- Content hashes of declared inputs form the cache key.
- Sandboxes or strict env filtering make declared inputs the real inputs.
- Finer graphs rerun less after an edit to shared code.
Alternative
Cache scripts that read whatever they like
Fast until an undeclared env var, host compiler or sibling file differs, and then confidently wrong.
- Wrong cache hits ship stale or mismatched outputs.
- Undeclared outputs go missing on a hit.
- Package-level keys churn whenever shared code changes.
From an edit to a trustworthy build
The hub map. Sibling pages cover graphs, cache keys, hermeticity, monorepos and resolution.
- 1
Read the graph
Targets, inputs and commands define what depends on what. - 2
Hash declared inputs
The cache key is a hash of every declared input of a task. - 3
Hit or run
Restore outputs on a hit, otherwise run the task, ideally in a sandbox. - 4
Catch undeclared reads
A task that reads hidden state makes later hits unsafe. Declare it, sandbox it, or mark it uncacheable.
The landscape: three different jobs
People lump these tools together, but they do three different jobs, and most real repos use one from each row.
| Category | Job | Unit of work | Examples | Knows about file-level inputs? |
|---|---|---|---|---|
| Task runner / monorepo orchestrator | Run existing scripts in graph order, skip repeats via cache | A package's task (web:build) | Make (as a runner), npm scripts, Turborepo, Nx, Lerna, Moon | Partly: hashes a package's files, not what each script reads |
| Build system | Own the build actions themselves, from fine-grained declared inputs | A target or action (//web:bundle, one compile) | Bazel, Buck2, Pants, Gradle, Make (as a file builder), Ninja, Please | Yes, every action's inputs and outputs are declared or inferred |
| Package / environment manager | Decide which versions of tools and libraries exist and make them available | A package or environment | npm, pnpm, Yarn, Cargo, Go modules, pip/uv, Nix, Guix, Docker, devcontainers, mise, asdf | It manages inputs for the other two rows |
The categories blur at the edges. Nx and Turborepo grew graph-aware caching that feels like a build system. Nix is a package manager whose store is literally a build system (the paper models it as one). Gradle is a build system with its own dependency resolver. The useful question is not "which category" but "what is the unit of work, and how precisely are its inputs known?"
What Turborepo, Nx, Bazel and Nix share, and where they differ
| Concept | Turborepo | Nx | Bazel | Nix |
|---|---|---|---|---|
| Shared idea | Hash inputs, look up a cache, skip identical work | Same | Same | Same |
| Unit of work | Package script (task) | Project target (task) | Action inside a target | Derivation (a package build step) |
| How the graph is known | package.json dependencies plus turbo.json dependsOn | Project graph from package files, imports and plugins | BUILD files declare every target and its deps | Nix expressions declare every input of a derivation |
| Cache key inputs | Package files, lockfile entries, task config, declared env vars, dependency task hashes | Project files, dependency files, external versions, runtime values, args | Action command line, input file digests, toolchain, platform | Hash of the whole derivation (builder, args, env, input store paths) |
| Who runs the actual compiler | Your existing scripts (tsc, next build) | Your scripts or Nx plugins/executors | Bazel rules, as sandboxed actions | Arbitrary builder in a sandbox |
| Hermeticity | Strict env mode filters env vars; no filesystem sandbox | No local filesystem sandbox; correctness relies on configured inputs | Per-action sandboxing, hermetic toolchains | Sandbox on by default on Linux, no network except fixed-output fetches |
| Remote | Remote cache (Vercel or self-hosted API) | Remote cache plus distributed task execution (Nx Cloud) | Remote cache and remote execution (Remote Execution API) | Binary caches (substituters) and remote builders |
| Adoption cost | Low: wrap existing scripts | Low to medium | High: rewrite builds as rules | Medium to high: new language and model |
| Sweet spot | JS/TS monorepos that want speed fast | JS/TS and polyglot monorepos wanting graph tooling and boundaries | Large polyglot repos needing correctness at scale | Reproducible toolchains, dev shells, OS images |
Notice the gradient: the further right you go, the more the tool insists on knowing every input, and the more it can safely cache and distribute in return.
Choosing: a decision chart
Decisions
- 1
Step 1: what hurts today?
- nextStep 2: slow repeated builds in a JS or TS monorepo?
- ?
Step 2: slow repeated builds in a JS or TS monorepo?
- nextStep 3: also need boundaries, generators, polyglot graph?
- nextStep 4: many languages, huge repo, wrong cache hits unacceptable?
- ?
Step 3: also need boundaries, generators, polyglot graph?
- nextPackage-level task runner with cache (Turborepo-style)
- nextGraph-aware orchestrator (Nx-style)
- 4
Package-level task runner with cache (Turborepo-style)
- 5
Graph-aware orchestrator (Nx-style)
- ?
Step 4: many languages, huge repo, wrong cache hits unacceptable?
- nextHermetic build system with remote execution (Bazel, Buck2, Pants)
- nextStep 5: problem is 'works on my machine' toolchains?
- 7
Hermetic build system with remote execution (Bazel, Buck2, Pants)
- ?
Step 5: problem is 'works on my machine' toolchains?
- nextEnvironment manager (Nix shell, devcontainer, mise) plus lockfiles
- nextKeep the simple tool, add a lockfile and CI cache first
- 9
Environment manager (Nix shell, devcontainer, mise) plus lockfiles
- 10
Keep the simple tool, add a lockfile and CI cache first
Runnable: the four properties in one tiny engine
The engine below builds app from two object files, then proves incrementality (only changed tasks re-run), caching (a hash of declared inputs is the key), correctness (the result matches a clean rebuild) and hermeticity (a task that peeks at CFLAGS without declaring it is refused).
"""A build system is a function from inputs to outputs over a graph.
This toy engine checks the four properties the hub page talks about:
1. correctness - every output equals what a clean rebuild would produce
2. incrementality - only tasks whose inputs changed are re-run
3. caching - results are looked up by a hash of declared inputs
4. hermeticity - a task may only read what it declared
Everything is in memory so it runs anywhere with plain Python 3.
"""
import hashlib, json
def h(obj) -> str:
"""Stable short hash of any JSON-serialisable value."""
return hashlib.sha256(json.dumps(obj, sort_keys=True).encode()).hexdigest()[:10]
# Source files: the "inputs" of the whole function.
src = {"util.c": "int add(int a,int b){return a+b;}",
"main.c": "int main(){return add(1,2);}"}
# Each task declares its inputs and a pure function that computes its output.
tasks = {
"util.o": (["util.c"], lambda r: "obj(" + r("util.c") + ")"),
"main.o": (["main.c"], lambda r: "obj(" + r("main.c") + ")"),
"app": (["util.o", "main.o"], lambda r: "link(" + r("util.o") + "+" + r("main.o") + ")"),
}
cache = {} # cache key -> output (a content-addressed action cache)
runs = [] # which tasks actually executed in the current build
def build(target, store, env, hermetic=True):
"""Build target, returning its value; store holds sources and built outputs."""
if target in src:
return store[target]
deps, fn = tasks[target]
dep_vals = [build(d, store, env, hermetic) for d in deps]
key = h([target, deps, dep_vals]) # declared inputs only
if key in cache: # property 3: caching
store[target] = cache[key]
return store[target]
def read(name):
# property 4: hermeticity - reading an undeclared name is an error
if hermetic and name not in deps:
raise PermissionError(f"{target} read undeclared input {name}")
return store[name] if name in store else env[name]
runs.append(target) # property 2: incrementality
store[target] = cache[key] = fn(read)
return store[target]
def clean_rebuild(target):
"""Reference answer: rebuild from scratch with no cache."""
store = dict(src)
def go(t):
if t in src: return store[t]
deps, fn = tasks[t]
for d in deps: go(d)
store[t] = fn(lambda n: store[n]); return store[t]
return go(target)
store = dict(src)
build("app", store, env={})
print("build 1 ran:", runs)
runs.clear(); build("app", store, env={})
print("build 2 (no change) ran:", runs)
src["main.c"] = "int main(){return add(2,3);}" # edit one file
store["main.c"] = src["main.c"]
runs.clear(); build("app", store, env={})
print("build 3 (main.c edited) ran:", runs)
print("property 1, correct vs clean rebuild:", store["app"] == clean_rebuild("app"))
# A task that sneaks a look at the environment breaks hermeticity.
tasks["main.o"] = (["main.c"], lambda r: "obj(" + r("main.c") + r("CFLAGS") + ")")
cache.clear(); store = dict(src)
try:
build("app", store, env={"CFLAGS": "-O2"})
except PermissionError as e:
print("hermetic engine refused:", e)Output:
build 1 ran: ['util.o', 'main.o', 'app']
build 2 (no change) ran: []
build 3 (main.c edited) ran: ['main.o', 'app']
property 1, correct vs clean rebuild: True
hermetic engine refused: main.o read undeclared input CFLAGSRunnable: the same edit at three granularities
Granularity is the biggest practical difference between a monorepo task runner and a build system. A package-level graph (Turborepo, Nx) treats a package as one unit, so editing a shared utility re-runs everything downstream. A target-level graph (Bazel, Buck2, Pants) knows which files import which, so it reruns far less, but only if dependencies are declared or inferred precisely.
// Same edit, three granularities: what gets re-run?
// - task runner, no cache: reruns every task (the Make-less npm scripts world)
// - package-level graph (Turborepo / Nx style): reruns the changed package and its dependents
// - target/file-level graph (Bazel / Buck2 / Pants style): reruns only targets whose inputs changed
// Repo: app -> ui -> utils, app -> api-client -> utils. Each package has a few source files.
type Pkg = { name: string; deps: string[]; files: string[] };
const pkgs: Pkg[] = [
{ name: "utils", deps: [], files: ["utils/date.ts", "utils/money.ts", "utils/strings.ts"] },
{ name: "ui", deps: ["utils"], files: ["ui/button.tsx", "ui/table.tsx"] },
{ name: "api-client", deps: ["utils"], files: ["api-client/http.ts", "api-client/orders.ts"] },
{ name: "app", deps: ["ui", "api-client"], files: ["app/page.tsx", "app/checkout.tsx"] },
];
// Fine-grained targets: which source files actually import which.
// Only orders.ts and checkout.tsx use utils/money.ts.
const fileDeps: Record<string, string[]> = {
"api-client/orders.ts": ["utils/money.ts", "api-client/http.ts"],
"app/checkout.tsx": ["api-client/orders.ts", "utils/money.ts"],
"ui/table.tsx": ["utils/strings.ts"],
"app/page.tsx": ["ui/button.tsx", "ui/table.tsx"],
};
const tasksPerPkg = ["build", "test", "lint"];
function closure(seed: string[], edges: Record<string, string[]>): Set<string> {
// Reverse-dependency closure: everything that (transitively) depends on a seed.
const out = new Set<string>(seed);
let grew = true;
while (grew) {
grew = false;
for (const [node, ds] of Object.entries(edges)) {
if (!out.has(node) && ds.some((d) => out.has(d))) { out.add(node); grew = true; }
}
}
return out;
}
const pkgEdges = Object.fromEntries(pkgs.map((p) => [p.name, p.deps]));
for (const changed of ["utils/money.ts", "app/page.tsx"]) {
const owner = pkgs.find((p) => p.files.includes(changed))!.name;
const affected = closure([owner], pkgEdges); // package granularity
const dirty = closure([changed], fileDeps); // file/target granularity
console.log(`edit: ${changed}`);
console.log(` task runner, no graph or cache: ${pkgs.length * tasksPerPkg.length} tasks`);
console.log(` package graph: ${[...affected].join(", ")} -> ${affected.size * tasksPerPkg.length} tasks`);
console.log(` file/target graph: ${[...dirty].join(", ")} -> ${dirty.size} compile units`);
}
console.log("trade-off: finer graphs rerun less but need every dependency declared (or inferred) precisely");Output:
edit: utils/money.ts
task runner, no graph or cache: 12 tasks
package graph: utils, ui, api-client, app -> 12 tasks
file/target graph: utils/money.ts, api-client/orders.ts, app/checkout.tsx -> 3 compile units
edit: app/page.tsx
task runner, no graph or cache: 12 tasks
package graph: app -> 3 tasks
file/target graph: app/page.tsx -> 1 compile units
trade-off: finer graphs rerun less but need every dependency declared (or inferred) preciselyExpectededit: utils/money.ts task runner, no graph or cache: 12 tasks package graph: utils, ui, api-client, app -> 12 tasks file/target graph: utils/money.ts, api-client/orders.ts, app/checkout.tsx -> 3 compile units edit: app/page.tsx task runner, no graph or cache: 12 tasks package graph: app -> 3 tasks file/target graph: app/page.tsx -> 1 compile units trade-off: finer graphs rerun less but need every dependency declared (or inferred) precisely
Press Run. Snippets must be self-contained — no network, files, or native modules.
Caching softens the package-level cost: downstream tasks whose own inputs and dependency outputs end up identical can still hit. But the hash of an upstream package changes on any edit, so its dependents' keys change too, and package-level tools miss more than target-level ones on shared code.
What happens if you choose otherwise
- Bazel for a 10-package TypeScript repo: you pay for writing and maintaining
BUILDfiles and rules for every tool, and the team spends weeks on build plumbing to save minutes. Many teams start with a task runner and migrate only when correctness or scale demands it. - A task runner for a huge polyglot repo: package-level granularity reruns too much, and scripts read whatever they like, so remote caching is either slow (low hit rate) or unsafe (undeclared inputs).
- Docker as your "build system": layer caching is a linear chain, so any change invalidates every later layer. The paper notes Docker-style tasks only form a linear chain. It is great for packaging and environments, weak as an incremental build graph.
- Nix for application builds at file granularity: Nix works at package (derivation) granularity. It shines for toolchains and system closures, but it is not designed to recompile one changed file inside a large app. Pair it with a language build tool.
Pitfalls
- Treating cache hits as proof of correctness. A hit only means the key matched. If the key misses a real input, the hit is wrong.
- Undeclared outputs. A task runner can only restore files it was told about (Turborepo caches only declared
outputs). Forgetting one gives a "successful" cached build with missing files. - Volatile tasks (stamping the time or git SHA into every build) defeat caching everywhere downstream. Isolate stamping into a final, cheap step.
- Assuming mtime is truth.
touch, clock skew, checkouts and backup restores all change mtimes without changing content, or the reverse.
The cluster map
- Dependency Graphs & Incremental Builds - Task DAGs, Topological Scheduling, Content Hashing & Affected Detection: package, task and action graphs, topological waves, mtime vs hashes, early cutoff and affected detection.
- Content-Addressed Caching - Cache Keys, Remote Cache, Remote Execution & Poisoning: what goes into a cache key, remote cache vs remote execution, cache economics and poisoning.
- Hermeticity & Reproducible Builds - Sandboxes, Pinned Toolchains, the Nix Store & Bit-for-Bit Output: pinned vs hermetic vs reproducible, SOURCE_DATE_EPOCH, the Nix store, and Nix vs Docker vs Bazel sandboxes vs lockfiles.
- Monorepo vs Polyrepo - Workspaces, Package Boundaries, Versioning & CI Fan-out: workspaces, package boundaries and code owners, fixed vs independent versioning, and CI fan-out by scale.
- Dependency Resolution & Dev Environments - SemVer, SAT vs MVS, Lockfiles, Nix Shells & Devcontainers: SemVer, search vs MVS vs nesting, lockfiles and integrity, hoisting, and Nix shells vs devcontainers vs mise.
Interview Q&A
What is the difference between a task runner and a build system?
Answer
A task runner orders and caches coarse tasks (run this package's build script) using whatever inputs it can hash around the script. A build system owns the actions and knows each one's precise inputs and outputs, so it can rebuild, cache and distribute at fine granularity and guarantee correctness.
Why do hermeticity and caching go together?
Answer
A cache key is a hash of declared inputs. If a task can read undeclared state, two runs with the same key can produce different outputs, and the cache will serve the wrong one. Hermeticity makes the declared inputs the real inputs.
How does Nix relate to Bazel if one is a package manager and the other a build system?
Answer
Both hash a precise description of every input and store results by that hash. Bazel does it per action with fine-grained file inputs; Nix does it per derivation (a package-level build step) and stores results in /nix/store. Build Systems a la Carte classifies both, with Bazel using constructive traces and Nix deep constructive traces.
When is Turborepo enough, and when would you reach for Bazel?
Answer
Turborepo is enough when most work is JS/TS package scripts, package-level granularity is acceptable and you can declare env vars and outputs. Reach for Bazel (or Buck2/Pants) when the repo is large and polyglot, you need remote execution, or wrong cache hits are unacceptable and you need sandboxed, fully declared actions.
Define a minimal build system.
Answer
One that runs each task at most once per build and only when something it transitively depends on has changed since the previous build.
Why can a cache hit be wrong?
Answer
The key matched, but an input that was not part of the key, such as an env var, /usr/bin/gcc or the time of day, was different. The cache returns an output built from other inputs.
Why is Docker a weak incremental build graph?
Answer
Docker layers form a linear chain, so any change invalidates every later layer. It is great for packaging and environments, not for rebuilding one changed file.
Where should build stamping such as the git SHA or build time go?
Answer
Into a final, cheap step. Stamping inside early tasks changes every downstream cache key and defeats caching everywhere.
Check yourself
Pick one CI job you own. List every input it reads, including env vars, tool versions and files outside its folder. Mark which ones are in its cache key today, and name the first undeclared input you would fix.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: CI/CD Pipelines — Stages, Artifacts, Caching & Supply Chain, CI Performance — Caching, Parallelism & Flaky Jobs, Graphs — Topological Sort & DAGs — Kahn, DFS Finish Times & Cycles, Terraform — Resources, Providers & the Dependency Graph, Cache Invalidation — TTL vs Event-Driven vs Versioned Keys, Git — Everyday Commands, Rebase vs Merge & Safe History.