DevOps
Part 2 of 6 · GitGit Objects, Refs, Index & Working Tree — Mental Model
Everything in Git is an immutable content-addressed object or a mutable ref. The index is the staging area between the working tree and the next commit. Rebase, reset, and reflog stop feeling like magic once those three trees are obvious.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What are the four object types?
Answer
Blob, tree, commit, and annotated tag. A lightweight tag is only a ref, not an object.
L2
Does a blob store its filename?
Answer
No. The tree stores the name, mode, and object id. Renaming a file with the same bytes keeps the blob id.
L3
What does git add actually write?
Answer
A blob if the bytes are new, and an index entry that maps the path and mode to that blob id.
L4
Why does amend change the commit hash?
Answer
The commit object bytes changed. A new tree, message, or timestamp is a new object. The branch ref moves to it.
L5
Where does HEAD point on a fresh clone of main?
Answer
HEAD is a symbolic ref to refs/heads/main, and that branch points at the commit fetched as origin/main.
L6
What is detached HEAD, and what fails if you commit there?
Answer
HEAD points at a commit, not a branch. Those commits become dangling when you switch away, unless you create a branch. The reflog still lists them for a while.
L7
Why does Git not track empty directories?
Answer
Trees exist because files exist. An empty directory has no tree entry. A placeholder file is the usual convention when a directory must stay.
Failure modes
Believing a blob knows its path
Rename detection is a heuristic on trees. The blob id is only the bytes.
Leaving an experiment in detached HEAD
Switching away without a new branch drops the only name for those commits. Recover from the reflog while it lasts.
Staging a whole file of debug and secrets
git add -p exists so the index can be a smaller snapshot than the working tree.
Hard reset of a dirty working tree
Uncommitted files are not objects yet. The object database cannot give them back.
Misconceptions
The index is a second copy of the working tree.
The index is the next commit. It can hold a subset of paths, and the working tree can be dirtier than that.
A branch is a chain of commits stored under the branch name.
A branch is a movable ref. The chain is parent pointers on commit objects.
Packed refs are a different kind of pointer.
packed-refs is a performance packing of the same ref names and ids.
Interviewer traps
Saying a rename always changes the blob hash.
Same bytes, same blob. The tree entry changes the name.
Describing rebase before you can say what a commit object contains.
Start with tree, parents, and the ref that moves. Rebase is the next page.
Design scenario
Same prompt for every reader.
Requirements
Show which objects changed, which ref moved, and how to keep the experiment commit.
Traffic / scale
One developer, one repository, a handful of files.
Latency
Local object writes. No network until fetch or push.
Consistency
The commit id is the hash of the commit object. Identical content and parents produce the same id.
Availability
A detached commit stays recoverable until the reflog expires, if no branch name was created.
Failure assumptions
- The developer switches back to main without creating a branch.
- The staged hunk is not the whole dirty file.
Constraints
- Do not claim the blob stores the filename.
- Keep the experiment by attaching a branch before leaving detached HEAD.
Prompt
Explain a repository where a developer renamed a file, staged one hunk, and then committed while HEAD was detached.
API
Which commands inspect HEAD, the index, and the commit object?
Data
What is inside the commit, the tree, and the blob after the rename?
Architecture
Where do refs/heads, packed-refs, and the index sit relative to the object database?
You renamed a file and want one clean commit
Prefer
Stage the hunk, then commit the index
The blob stays stable when the bytes stay stable. The tree records the new name. The index holds only the hunk you meant to record.
- git add -p keeps debug and secrets out of the commit.
- git diff --staged is the review of the index against HEAD.
- The commit id changes only because the tree or metadata changed.
Alternative
Commit the whole dirty working tree
Every untracked experiment and secret in the file becomes part of the snapshot.
- The blob id changes because the bytes changed, not because of the rename.
- You cannot unstage one path later without another command.
- Reviewers see noise instead of the rename.
From a file on disk to a commit object
Vertical cards for the same path as the diagram. Nothing here rewrites a shared branch.
- 1
Edit the working tree
The file on disk can be dirty. Git has not stored those bytes as a new blob yet. - 2
git add
Write a blob if needed. Point the index entry for that path at the blob id and mode. - 3
git commit
Write a tree from the index and a commit that points at the tree and at the parent. - 4
Move the branch ref
The branch now names the new commit. HEAD follows the branch, unless you are detached.
Overview
Once you can point at an object or a ref, rebase and reset are pointer motion plus new objects. They are not a second database.
Interviews start here even when the question sounds like a workflow question. If you cannot say what a commit contains, the later pages are slogans.
Object types
| Type | Payload | When created | Why it beats the alternative |
|---|---|---|---|
| blob | File bytes, no filename | add or commit of new content | Identical files share one object across trees |
| tree | Mode, name, and object id per entry | commit | Directory structure stays out of the blob |
| commit | Tree, parents, author, message | commit, merge, cherry-pick | A snapshot id for bisect and revert |
| annotated tag | Object, message, optional signature | tag -a | A release label with metadata, unlike a lightweight ref |
Failure if you confuse them: a blob does not know its filename. The tree does. Rename reasoning breaks if you treat the blob as a path.
Three trees
| Area | Inspect | Mutable? | Risk |
|---|---|---|---|
| Working tree | git status, git diff | Yes | Uncommitted loss from hard reset or clean |
| Index | git diff --staged, git ls-files -s | Yes | The wrong hunks are staged |
| HEAD commit | git show, git cat-file -p HEAD | Only by new commits or reset | Rewriting a shared tip |
Flow
- 1
1. File on disk
- next2. git add
- 2
2. git add
- next3. Index path to blob
- 3
3. Index path to blob
- next4. git commit
- 4
4. git commit
- next5. Tree
- 5
5. Tree
- next6. Blob bytes
- 6
6. Blob bytes
Lesson map
Git Objects, Refs, Index & Working Tree — Mental Model
Everything in Git is an immutable content-addressed object or a mutable ref. The index is the staging area between the working tree and the next commit. Rebase, reset, and reflog stop feeling like magic once those three trees are obvious.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB file["1. File on disk"] add["2. git add"] idx["3. Index path to blob"] cmt["4. git commit"] file -->|1. File on disk to 2. git add| add add -->|2. git add to 3. Index path to blob| idx idx -->|3. Index path to blob| cmt
Refs and HEAD
git rev-parse HEAD
git rev-parse main
git symbolic-ref HEAD
git show-ref --heads --tags
git log -1 --format='%H %D'| Ref kind | Example | Moves when |
|---|---|---|
| Branch | refs/heads/feature | commit, reset, rebase, or merge on that branch |
| Remote-tracking | refs/remotes/origin/main | fetch |
| Lightweight tag | refs/tags/v1 | Rarely. Treat published tags as immutable |
| HEAD | Symbolic, or detached | checkout or switch |
Flow
- 1
1. HEAD symbolic
- next2. refs/heads/feature
- 2
2. refs/heads/feature
- next3. commit C3
- 3
3. commit C3
- next4. commit C2
- 4
4. commit C2
- next5. commit C1
- 5
5. commit C1
Detached HEAD
git switch --detach v1.2.0
git switch -c keep-experiment
git switch mainThe first command points HEAD at a commit. The second attaches a branch so the experiment stays reachable. Switching to main without that branch leaves the experiment commits dangling. They remain in the reflog for a while. The recovery lesson is where you go looking.
Staging patterns
| Goal | Prefer | Beats | Failure if wrong |
|---|---|---|---|
| Partial commit | git add -p | Staging a whole file of cruft | Secrets in the commit |
| Unstage, keep edits | git restore --staged PATH | Muscle memory that discards the working tree | Accidental loss of edits |
| Discard a working-tree file | git restore PATH | A mix of delete and checkout | Losing the only copy |
| See the exact index | git ls-files -s | Guessing | Missing mode 100755 versus 100644 |
git status -sb
git add -p path/to/file.go
git diff --staged
git commit -m "fix: nil guard on checkout"
git restore --staged path/to/file.goPlumbing peek
You rarely need plumbing on a normal day. Knowing it exists is the interview proof that the model is real.
git cat-file -t HEAD
git cat-file -p HEAD
git rev-parse 'HEAD^{tree}'
git ls-tree -r HEADgit hash-object can write a blob. Use it in a scratch repository, not to smuggle bytes into a shared repo by accident.
What the hash actually covers
A commit id changes when the tree, a parent, the message, or the author metadata changes. A rebase copies commits, so the parents change, so the ids change. Amend does the same to the tip. The branch ref is the only thing that "moves" in the way people mean. The old commit object stays in the object database until nothing, including the reflog, points at it.
Empty directories are not tracked. Git stores trees because files exist. A .gitkeep file is a convention, not a special object type.
Same bytes, new name
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Does renaming a file change the blob id?
Answer
Not if the bytes are unchanged. The tree entry changes the name. Status may show a rename as a heuristic. The blob hash stays put.
Why does amending change the commit hash?
Answer
The commit object bytes changed. A new tree, message, or timestamp is a new object. The branch ref now points at that object.
What does git add actually do?
Answer
It writes a blob when the content is new and updates the index entry for that path to the new object id and mode.
How is a branch different from a tag?
Answer
A branch is expected to move. An annotated tag marks a point, usually a release, and can carry a message and a signature. Do not move a published tag casually.
Where is HEAD after a fresh clone on main?
Answer
HEAD is a symbolic ref to refs/heads/main. That branch points at the commit you fetched as origin/main.
Can the index and the working tree both differ from HEAD?
Answer
Yes. That is the normal dirty state. Some tracked files are modified, a subset is staged, and HEAD is still the last commit.
Why are empty directories missing?
Answer
Git tracks trees through file paths. An empty directory has nothing to hash. Add a placeholder file if the directory must exist.
What are packed refs?
Answer
Many refs stored in one packed-refs file for performance. The names and the commit ids mean the same thing as loose ref files.
Pitfalls
- Drawing a branch as a folder of commits. It is one file, or one line in packed-refs, holding one id.
- Using
resetwhen you meantrestore --staged, and wiping the working tree. - Committing in detached HEAD during a bisect and then running
bisect resetwithout saving a branch. - Treating a lightweight tag and an annotated tag as the same object. One is a ref. The other is an object plus a ref.
Pick a file, change only its path, and say which object id stays the same. Then stage one hunk of a second edit and say which of the three trees changed.