The short answer
Quick answer: Git is a content-addressed database of snapshots. Every version of every file is stored as a blob, every directory as a tree listing its blobs and subtrees, and every commit as a small object pointing to one tree plus its parent commit(s). Each object is named by the hash of its contents, so identical content is stored once and nothing can change without its name changing. A branch is just a tiny file containing the hash of a commit. Commits link to their parents, forming a graph that is your project's history.
Snapshots, not differences
A common assumption is that Git stores a list of changes. Conceptually, it does not. Each commit records the complete state of the project at that moment.
That sounds wasteful, but files that did not change are not stored again. The new snapshot simply points to the same objects as the old one. This is the structural sharing idea described in why immutability matters.
The object database
Everything lives in the .git/objects directory. There are four kinds of object, explained in detail in the Pro Git chapter on Git objects:
| Object | Contains | Analogy |
|---|---|---|
| Blob | The contents of one file (no name) | A file's data |
| Tree | A list of names, modes, and the hashes of blobs or other trees | A directory |
| Commit | A tree hash, parent hash(es), author, date and message | A snapshot with context |
| Tag | A named, optionally signed pointer to a commit | A label |
Each object's name is a cryptographic hash of its contents (a much stronger kind of hash than the lightweight ones used in hash maps): SHA-1 by default, a 40-character hexadecimal string, with SHA-256 available in newer versions. This has three important consequences:
- Deduplication. Two identical files, anywhere in history, are the same blob.
- Integrity. If a single bit of an object is corrupted, its hash no longer matches.
- Tamper evidence. A commit's hash covers its tree and its parents, so it effectively covers the whole history behind it. You cannot alter an old commit without changing the hash of every commit after it.
You can look inside any object:
git cat-file -t HEAD # type: commit
git cat-file -p HEAD # show the commit object
git cat-file -p HEAD^{tree} # show its root tree
How a commit is built
Suppose you edit README.md in a project with two files and commit:
- Git creates a new blob for the new contents of
README.md. - It creates a new tree for the root directory, listing the new blob and the existing, unchanged blob for the other file.
- It creates a commit pointing to that tree, with the previous commit as its parent.
Unchanged files and directories cost nothing. Only the changed blob and the trees on the path up to the root are new.
Because each commit names its parent, commits form a directed acyclic graph. A normal commit has one parent; a merge commit has two or more; the first commit has none.
Branches are just pointers
A branch is a file in .git/refs/heads/ containing one commit hash. That is the whole thing.
- Creating a branch writes 41 bytes. This is why branching in Git is instant.
- Committing creates the new commit and moves the current branch pointer to it.
HEADis a pointer to the branch you are on (or directly to a commit, a state called "detached HEAD").- Tags are pointers that do not move.
- Remote-tracking branches such as
origin/mainrecord where the remote's branches were when you last fetched.
Deleting a branch deletes the pointer, not the commits.
The three areas
| Area | What it is |
|---|---|
| Working directory | The actual files you edit |
| Index (staging area) | The proposed contents of the next commit |
| Repository | The object database and refs |
git add writes the file's contents as a blob and records it in the index. git commit turns the index into tree objects and creates the commit. The staging area lets you build a commit from only some of your changes.
Merge and rebase
Merging two branches finds their most recent common ancestor and combines the changes each side made since then (a three-way merge).
- If your branch has not diverged, Git simply moves the pointer forward: a fast-forward.
- Otherwise it creates a merge commit with two parents.
- A conflict arises when both sides changed the same lines, and you must choose.
Rebasing replays your commits one at a time on top of another commit. Since a commit's hash depends on its parent, the replayed commits are new commits with new hashes. The old ones are abandoned. This is why you should not rebase commits that other people have already based work on.
| Merge | Rebase | |
|---|---|---|
| History | Preserved as it happened | Rewritten to be linear |
| New commits | One merge commit | A new copy of each of your commits |
| Safe on shared branches | Yes | No |
Packfiles: where the space savings come from
Storing a full copy of a large file for every small edit would add up. So Git periodically bundles objects into packfiles. Inside a pack, similar objects are stored as one full copy plus compact deltas describing the differences, and everything is compressed. The Pro Git chapter on packfiles shows this in action.
So deltas are a storage optimisation underneath; the model you work with is still whole snapshots. Packs are also what gets sent over the network during fetch and push.
It is hard to lose work
Objects are immutable, and commands such as commit --amend, rebase and reset create new objects or move pointers; they do not destroy the old commits straight away.
The reflog records every position HEAD has had:
git reflog
git reset --hard HEAD@{2} # go back to where you were two moves ago
Unreachable objects are only removed by garbage collection, typically after a grace period of weeks. If you committed it, you can almost always get it back. Uncommitted changes are the exception: Git cannot recover what it never stored.
Distributed by design
Every clone contains the full object database and history. Fetching and pushing are simply "send me the objects I do not have, then update this pointer". Because objects are identified by their content, two repositories agree automatically on what any given hash means. Platforms such as GitHub add collaboration features on top, including the automation described in how CI/CD pipelines work.
Frequently asked questions
Does Git store diffs or snapshots?
Snapshots. Each commit points to a complete tree. Diffs are computed when you ask for them, and deltas are used only as a compression technique inside packfiles.
What is a commit hash?
The hash of the commit object's contents: its tree, parent hashes, author, date and message. Change any of them and the hash changes.
Why is creating a branch so fast?
Because a branch is a single small file containing a commit hash. Nothing is copied.
What does "detached HEAD" mean?
HEAD points directly at a commit instead of at a branch. New commits you make are not on any branch, so create one before switching away.
Conclusion
Git's design is small and elegant: immutable objects named by their hash, linked into a graph, with branches as movable labels. Once you picture commands as "create objects" or "move pointers", operations like merge, rebase and reset stop being incantations and start making sense.
Related articles
- Why Immutability Makes Code Easier to Reason About
- How Hash Maps Achieve O(1) Lookups
- How File Systems Store Your Files on Disk
- How CI/CD Pipelines Ship Code Automatically
