All notes

Git Doesn't Compare Every File. Here's What It Actually Does.

Running git status on a huge repository feels instant. Surely Git isn't scanning millions of files every time... right?

4 min read

Every developer has typed this command.

git status

A second later, Git somehow knows:

Modified:
  UserService.java
 
Deleted:
  config.yml
 
Untracked:
  notes.md

It almost feels like magic.

Did Git compare every file in your repository?

Not quite.

Let's see what actually happens.

The Naive Approach

Imagine a repository with:

2,000,000 files

A naive algorithm would do:

For every file
 
↓
 
Open file
 
↓
 
Read entire contents
 
↓
 
Compare with previous version

That would make every git status painfully slow.

Instead, Git is much smarter.

The Git Index

Git maintains something called the Index (also called the Staging Area).

Think of it as a snapshot of your working tree.

Working Directory
 
↓
 
Git Index
 
↓
 
Last Commit

The index stores metadata about every tracked file.

For example:

README.md
 
Size:
2048 bytes
 
Modified:
10:42:11
 
Hash:
abc123...

Git doesn't immediately read the file contents.

It first compares inexpensive metadata.

Fast Path

Suppose you never touched a file.

Its:

are still identical.

Git assumes:

Nothing changed.

No file reading.

No hashing.

No expensive work.

Slow Path

Now imagine you edit:

UserService.java

Its modified time changes.

Git notices the metadata no longer matches.

Only then does it read the file and compute a new SHA-1 (or SHA-256 in newer repositories).

Old Hash
 
↓
 
abc123
 
↓
 
New Hash
 
↓
 
f91d44

If the hashes differ...

Git marks the file as modified.

Why Hashes?

Git identifies file contents using cryptographic hashes.

Two identical files produce the same hash.

Hello World
 
↓
 
SHA
 
↓
 
7b502...

Change even one character.

Hello World!
 
↓
 
SHA
 
↓
 
91fae...

Completely different hash.

This makes detecting content changes extremely reliable.

What Happens During a Commit?

When you run:

git commit

Git doesn't save a list of changes.

It creates immutable objects.

Blob
 
↓
 
Tree
 
↓
 
Commit

Blob

Stores the file contents.

Tree

Represents folders.

Contains references to blobs.

Commit

Points to a tree and its parent commit.

The result is a directed graph of snapshots rather than a sequence of patches.

Why Cloning Is Surprisingly Efficient

Imagine two commits.

Commit A
 
↓
 
README.md

and

Commit B
 
↓
 
README.md

If the file never changed...

Git doesn't duplicate it.

Both commits simply reference the same blob object.

Commit A
 
      \
 
       Blob
 
      /
 
Commit B

One copy.

Multiple references.

This saves enormous amounts of storage.

Why Git Is So Fast

Git optimizes almost every operation.

Instead of repeatedly reading every file, it uses:

Only the files that actually changed require expensive work.

Interview Questions

Why doesn't Git compare every file?

Because checking metadata is much cheaper than reading file contents.

Git only hashes files whose metadata indicates they may have changed.


Why use hashes?

Hashes uniquely identify file contents.

If two hashes differ, Git knows the contents differ.


Does Git store diffs?

Not internally.

Git stores snapshots of files as immutable objects.

Diffs are generated when needed.


What is the staging area?

It's Git's index.

A snapshot that sits between your working directory and the next commit.

It allows you to choose exactly what goes into a commit.

Final Thoughts

The next time you run:

git status

remember that Git isn't opening every file in your repository.

It first asks a much cheaper question:

"Has anything about this file changed enough to make reading it worthwhile?"

That small optimization is one of the reasons Git remains incredibly fast, even for repositories containing millions of files.