Most Git trouble isn't a memory problem. People know the commands, but they can't predict what a command will do, so the first unexpected output sends them searching for a scary-sounding fix. Once you understand how Git works underneath — three areas where your files live, a small set of object types, and branches that are nothing more than labels — most of that unpredictability goes away. This is Part 1 of a series on using Git in real projects, and it stays hands-on. Every command below was run in a scratch repository on Git 2.43, and you can repeat all of it in about five minutes.
Set up a scratch repository
You need Git installed and a throwaway directory. If you have never committed before, set your identity first, or git commit will refuse to run:
git config --global user.name "Your Name"
git config --global user.email "you@example.com"
mkdir git-demo && cd git-demo
git init -b main
echo "hello git" > notes.txt
The -b main flag names the first branch. Without it, git init currently creates master; the Git docs say the default changes to main in Git 3.0. You can set it once for every new repository with git config --global init.defaultBranch main.
The three areas
Git keeps your content in three places, and nearly every command moves content between them:
| Area | What it is | What it holds |
|---|---|---|
| Working tree | The actual files in your project folder | What you edit |
| Index (staging area) | A file at .git/index
|
The proposed next commit |
| Repository | Objects and references inside .git
|
Every commit ever made, plus where HEAD points |
The Pro Git book describes the index as your "proposed next commit," and that phrase is the most useful one to remember. When you run git commit, Git does not look at your working files at all. It records whatever is in the index.
Watch the file move through the areas with git status -s, which prints two status columns. The first compares the index to the last commit, and the second compares the working tree to the index:
git status -s
?? notes.txt # untracked: exists only in the working tree
git add notes.txt
git status -s
A notes.txt # A in the first column: staged as a new file
git ls-files -s
100644 8d0e41234f24b6da002d962a26c2495ea16a425f 0 notes.txt
git ls-files -s prints what is in the index: the file mode, a hash, and the path. That hash identifies the file's content, and the content was captured at the moment you ran git add. If you edit notes.txt again, git status -s shows MM: one version staged, a different version in the working tree. Committing would record the staged one.
git commit -m "Add notes"
git log --oneline --decorate
a689b02 (HEAD -> main) Add notes
Your commit hash will differ from a689b02, because a commit hash covers the author, the timestamp and the message as well as the content.
Three flavors of git diff
Once you see the three areas, the variants of git diff stop being arbitrary. Each one compares a different pair:
| Command | Compares | Answers |
|---|---|---|
git diff |
Working tree against index | What have I changed but not staged? |
git diff --staged |
Index against last commit | What will my next commit contain? |
git diff HEAD |
Working tree against last commit | What is different from the last commit overall? |
What a commit actually stores
Git's database holds a few kinds of objects. Three matter for everyday work:
- A blob stores the contents of one file. It stores no filename.
- A tree stores a directory listing: names, file modes, and the hashes of the blobs and sub-trees inside.
- A commit points to one tree (the whole project at that moment), to its parent commit or commits, and carries the author, the committer and the message.
A fourth type, the annotated tag, comes up when we get to releases in Part 9. You can inspect any object with git cat-file: -t prints its type and -p prints its contents.
git cat-file -p HEAD
tree 87ee75ad134a458a58b7d048c63f6e88dea893ac
author Your Name <you@example.com> 1791380841 +0530
committer Your Name <you@example.com> 1791380841 +0530
Add notes
git cat-file -p 'HEAD^{tree}'
100644 blob 8d0e41234f24b6da002d962a26c2495ea16a425f notes.txt
git cat-file -t HEAD
commit
If you have commit signing turned on, the commit will show an extra gpgsig block; everything else looks the same. Follow the chain: the commit names a tree, and the tree names the blob. The blob and tree hashes above are the same on any machine, because they depend only on the content. If you created notes.txt with exactly hello git and a trailing newline, you will get these identical hashes.
Same content, same hash
Git names every object by a hash of its content. You can check that by hashing the same text yourself:
echo "hello git" | git hash-object --stdin
8d0e41234f24b6da002d962a26c2495ea16a425f
git rev-parse HEAD:notes.txt
8d0e41234f24b6da002d962a26c2495ea16a425f
The two hashes match. Per the Pro Git book, Git computes the hash over a small header (such as blob <size> followed by a null byte) plus the content, compresses the result, and stores it under .git/objects using the first two hex characters as a directory name:
find .git/objects -type f | sort
.git/objects/87/ee75ad134a458a58b7d048c63f6e88dea893ac # the tree
.git/objects/8d/0e41234f24b6da002d962a26c2495ea16a425f # the blob
.git/objects/a6/89b02d158ae9eb63740f7ab220d6254b0193e8 # the commit
This design has practical consequences:
- A file that doesn't change between commits is not stored again. Both commits point to the same blob.
- Logically, each commit is a full snapshot of the project, not a diff. (Git does compress stored objects against each other in packfiles to save space, but that is a storage detail and not how you should think about history.)
- Git tracks files, not folders. A tree is built from files, so an empty directory has nothing to record. Projects that need an empty folder usually add a placeholder file such as
.gitkeep. - The hash length you see depends on the repository's hash function. Repositories on Git 2.43 default to 40-character SHA-1 hashes, which is what this series shows. Git also documents an option to create SHA-256 repositories, but that is not what you will meet in most projects today.
Branches are just pointers
A branch is a file that contains one commit hash. Look at it:
cat .git/refs/heads/main
a689b02d158ae9eb63740f7ab220d6254b0193e8
wc -c .git/refs/heads/main
41 .git/refs/heads/main
Forty characters plus a newline. That is the whole branch. (Git can later pack references into a single .git/packed-refs file, for example after git gc, so you may not always find the individual file. The idea doesn't change.) This is why creating a branch is instant and costs nothing, regardless of how big the project is:
git branch feature
git log --oneline --decorate
a689b02 (HEAD -> main, feature) Add notes
Both labels sit on the same commit. git branch only created a pointer; nothing was copied.
HEAD: where you are right now
HEAD tells Git which commit you are working from. Normally it doesn't point at a commit directly. It points at a branch, which points at a commit:
cat .git/HEAD
ref: refs/heads/main
git switch feature
cat .git/HEAD
ref: refs/heads/feature
Switching branches moves HEAD and also updates your working tree and index to match the snapshot of the branch you moved to. When you commit, the branch that HEAD points at moves forward and the others stay where they were:
echo "change" >> notes.txt
git commit -am "Edit notes on feature"
git switch main
git log --oneline --decorate --graph --all
* 3c4cdbb (feature) Edit notes on feature
* a689b02 (HEAD -> main) Add notes
Detached HEAD
You can also point HEAD straight at a commit, which Git calls a detached HEAD. You end up there when you check out a tag or a specific commit hash:
git switch --detach a689b02
git status
HEAD detached at a689b02
cat .git/HEAD
a689b02d158ae9eb63740f7ab220d6254b0193e8
This is fine for looking around. The risk is committing while detached: those commits aren't on any branch, so when you switch away nothing points to them anymore. If you want to keep work done in this state, give it a branch name first with git switch -c new-branch-name. To get back to normal, run git switch main.
The reflog: Git's local undo history
Git also keeps a log of where HEAD has been, so even commits that no branch points to can usually be found again:
git reflog
a689b02 HEAD@{0}: checkout: moving from feature to main
3c4cdbb HEAD@{1}: commit: Edit notes on feature
a689b02 HEAD@{2}: checkout: moving from main to feature
a689b02 HEAD@{3}: commit (initial): Add notes
HEAD@{2} means "where HEAD was two moves ago." Per the Git documentation, the reflog is local to your repository, and entries expire after 90 days by default (30 days for entries no longer reachable from any branch), controlled by gc.reflogExpire and gc.reflogExpireUnreachable. Part 5 uses it to recover lost commits and deleted branches. For now, it's enough to know it exists and that it's your safety net.
What common commands do to the three areas
| Command | What moves |
|---|---|
git add <file> |
Working tree → index |
git commit |
Index → new commit; the current branch moves to it |
git switch <branch> |
HEAD moves; index and working tree are updated to match |
git restore --staged <file> |
Last commit → index (unstages the file) |
git restore <file> |
Index → working tree (discards your unstaged edits) |
Be careful with that last row. Edits that you never staged were never recorded in Git's database, so git restore <file> throws them away and the reflog can't bring them back. Staging a file first, even without committing, gives Git a copy to work with.
Common misconceptions this model clears up
| What people think | What's actually happening |
|---|---|
"I edited the file after git add, and the edit was committed." |
The index holds the content from when you ran git add. Later edits are separate until you stage again. |
| "Deleting a branch deletes the commits." | It deletes a pointer. The commits still exist and can be found through the reflog until Git cleans them up. |
| "A branch is a copy of the code." | A branch is a 41-byte file naming one commit. |
| "Git stores diffs between versions." | Each commit is a full snapshot. Git saves space by reusing identical content and compressing storage. |
| "My empty folder isn't showing up in the repository." | Git tracks files. A folder only exists in a commit if it contains at least one tracked file. |
"git switch only changes the branch name shown." |
It also rewrites your working tree to match that branch's snapshot. |
Before you move on
Try to predict, before you run it, what git status and git diff --staged will show after each of your next few edits and git add commands. If your guess is wrong, work out which area the content was actually in. That habit is worth more than memorizing flags, and it is the foundation for the next parts: staging precisely in Part 2, then branching, merging and rebasing in Parts 3 and 4, where the idea that a branch is just a movable pointer does most of the explaining.
References
- Pro Git: Git Objects
- Pro Git: Branches in a Nutshell
- Pro Git: Reset Demystified (the three trees)
- Git documentation: git-reflog
- Git documentation: git-init
Next in the series: Part 2: Your Daily Git Workflow Done Right.
Originally published on DEV Talk.
Top comments (0)