Search

How Git Stores Your Code History

The short answer

Quick answer: Git is a content-addressed database of snapshots. Every version of every file is stored as a blob, every directory as a tree listing its blobs and subtrees, and every commit as a small object pointing to one tree plus its parent commit(s). Each object is named by the hash of its contents, so identical content is stored once and nothing can change without its name changing. A branch is just a tiny file containing the hash of a commit. Commits link to their parents, forming a graph that is your project's history.

Snapshots, not differences

A common assumption is that Git stores a list of changes. Conceptually, it does not. Each commit records the complete state of the project at that moment.

That sounds wasteful, but files that did not change are not stored again. The new snapshot simply points to the same objects as the old one. This is the structural sharing idea described in why immutability matters.

The object database

Everything lives in the .git/objects directory. There are four kinds of object, explained in detail in the Pro Git chapter on Git objects:

ObjectContainsAnalogy
BlobThe contents of one file (no name)A file's data
TreeA list of names, modes, and the hashes of blobs or other treesA directory
CommitA tree hash, parent hash(es), author, date and messageA snapshot with context
TagA named, optionally signed pointer to a commitA label

Each object's name is a cryptographic hash of its contents (a much stronger kind of hash than the lightweight ones used in hash maps): SHA-1 by default, a 40-character hexadecimal string, with SHA-256 available in newer versions. This has three important consequences:

  • Deduplication. Two identical files, anywhere in history, are the same blob.
  • Integrity. If a single bit of an object is corrupted, its hash no longer matches.
  • Tamper evidence. A commit's hash covers its tree and its parents, so it effectively covers the whole history behind it. You cannot alter an old commit without changing the hash of every commit after it.

You can look inside any object:

git cat-file -t HEAD            # type: commit
git cat-file -p HEAD            # show the commit object
git cat-file -p HEAD^{tree}     # show its root tree

How a commit is built

Suppose you edit README.md in a project with two files and commit:

  1. Git creates a new blob for the new contents of README.md.
  2. It creates a new tree for the root directory, listing the new blob and the existing, unchanged blob for the other file.
  3. It creates a commit pointing to that tree, with the previous commit as its parent.

Unchanged files and directories cost nothing. Only the changed blob and the trees on the path up to the root are new.

Because each commit names its parent, commits form a directed acyclic graph. A normal commit has one parent; a merge commit has two or more; the first commit has none.

Branches are just pointers

A branch is a file in .git/refs/heads/ containing one commit hash. That is the whole thing.

  • Creating a branch writes 41 bytes. This is why branching in Git is instant.
  • Committing creates the new commit and moves the current branch pointer to it.
  • HEAD is a pointer to the branch you are on (or directly to a commit, a state called "detached HEAD").
  • Tags are pointers that do not move.
  • Remote-tracking branches such as origin/main record where the remote's branches were when you last fetched.

Deleting a branch deletes the pointer, not the commits.

The three areas

AreaWhat it is
Working directoryThe actual files you edit
Index (staging area)The proposed contents of the next commit
RepositoryThe object database and refs

git add writes the file's contents as a blob and records it in the index. git commit turns the index into tree objects and creates the commit. The staging area lets you build a commit from only some of your changes.

Merge and rebase

Merging two branches finds their most recent common ancestor and combines the changes each side made since then (a three-way merge).

  • If your branch has not diverged, Git simply moves the pointer forward: a fast-forward.
  • Otherwise it creates a merge commit with two parents.
  • A conflict arises when both sides changed the same lines, and you must choose.

Rebasing replays your commits one at a time on top of another commit. Since a commit's hash depends on its parent, the replayed commits are new commits with new hashes. The old ones are abandoned. This is why you should not rebase commits that other people have already based work on.

MergeRebase
HistoryPreserved as it happenedRewritten to be linear
New commitsOne merge commitA new copy of each of your commits
Safe on shared branchesYesNo

Packfiles: where the space savings come from

Storing a full copy of a large file for every small edit would add up. So Git periodically bundles objects into packfiles. Inside a pack, similar objects are stored as one full copy plus compact deltas describing the differences, and everything is compressed. The Pro Git chapter on packfiles shows this in action.

So deltas are a storage optimisation underneath; the model you work with is still whole snapshots. Packs are also what gets sent over the network during fetch and push.

It is hard to lose work

Objects are immutable, and commands such as commit --amend, rebase and reset create new objects or move pointers; they do not destroy the old commits straight away.

The reflog records every position HEAD has had:

git reflog
git reset --hard HEAD@{2}     # go back to where you were two moves ago

Unreachable objects are only removed by garbage collection, typically after a grace period of weeks. If you committed it, you can almost always get it back. Uncommitted changes are the exception: Git cannot recover what it never stored.

Distributed by design

Every clone contains the full object database and history. Fetching and pushing are simply "send me the objects I do not have, then update this pointer". Because objects are identified by their content, two repositories agree automatically on what any given hash means. Platforms such as GitHub add collaboration features on top, including the automation described in how CI/CD pipelines work.

Frequently asked questions

Does Git store diffs or snapshots?

Snapshots. Each commit points to a complete tree. Diffs are computed when you ask for them, and deltas are used only as a compression technique inside packfiles.

What is a commit hash?

The hash of the commit object's contents: its tree, parent hashes, author, date and message. Change any of them and the hash changes.

Why is creating a branch so fast?

Because a branch is a single small file containing a commit hash. Nothing is copied.

What does "detached HEAD" mean?

HEAD points directly at a commit instead of at a branch. New commits you make are not on any branch, so create one before switching away.

Conclusion

Git's design is small and elegant: immutable objects named by their hash, linked into a graph, with branches as movable labels. Once you picture commands as "create objects" or "move pointers", operations like merge, rebase and reset stop being incantations and start making sense.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy