All notes

CAP Theorem Isn't What Most Engineers Think

CAP Theorem doesn't mean choosing any two out of three. The tradeoff appears when a network partition prevents distributed nodes from communicating.

7 min read


Most engineers learn CAP Theorem like this:

"Choose any two: Consistency, Availability, and Partition Tolerance."

It's a nice mnemonic.

It's also misleading.

CAP doesn't say you permanently have to choose two out of three.

The real idea is more specific:

When a network partition occurs, a distributed system cannot simultaneously guarantee both strong Consistency and Availability.

That's the actual tradeoff.


First, What Does CAP Mean?

CAP stands for:

Let's understand each one without the usual textbook definition.

Consistency

Every successful read receives the most recent successful write.

Imagine two replicas:

Server A
Balance = $100
 
Server B
Balance = $100

A user deposits $50.

After the write succeeds:

Server A
Balance = $150
 
Server B
Balance = $150

A strongly consistent system shouldn't allow another user to successfully read $100 from Server B after the system has acknowledged the $50 deposit.


Availability

Every request to a non-failing node receives a response.

The important word is:

response.

An available system doesn't necessarily promise that the response contains the newest data.

It promises that the system will respond.


Partition Tolerance

A partition happens when nodes in a distributed system can no longer communicate reliably.

For example:

              Network
                 X
                 X
                 X
 
          Server A       Server B
             |              |
             |              |
          Database        Database

Both servers are alive.

But they can't communicate.

That's a network partition.

And this is where CAP becomes interesting.


CAP Doesn't Matter When Everything Is Working

Suppose we have:

Server A ←──── Network ────→ Server B

Everything is communicating normally.

A write arrives at Server A.

Server A replicates it to Server B.

Both replicas stay synchronized.

You can have:

Consistency
     +
Availability
     +
Partition tolerance

There is no partition to force a tradeoff.

The CAP constraint becomes relevant when the partition actually happens.


Now Break The Network

Imagine two database replicas.

Client
  |
  v
Server A   X────────X   Server B

The connection between them is gone.

Now imagine a user updates their account on Server A:

Balance: $100
 
       ↓
 
Deposit $50
 
       ↓
 
Server A = $150
Server B = $100

Server A knows about the new balance.

Server B doesn't.

Now another user asks Server B:

"What's the balance?"

The system has a problem.

It has two competing choices.


Choice 1: Preserve Consistency

Server B could refuse to answer until it knows the latest state.

Client
  |
  v
Server B
  |
  X
"Cannot safely answer"

You sacrifice availability.

The system says:

"I'd rather fail this request than potentially return incorrect or stale information."

This is the CP side of the tradeoff.

Consistency ✓
Partition Tolerance ✓
Availability ✕

Choice 2: Preserve Availability

Server B could immediately return its local value:

$100

The request succeeds.

The system remains available.

But the value may be stale.

Consistency ✕
Partition Tolerance ✓
Availability ✓

This is the AP side.

The system chooses to remain responsive even though replicas may temporarily disagree.


So What Does "CP vs AP" Actually Mean?

It doesn't mean:

"This database is permanently CP."

or:

"That database is permanently AP."

The more useful way to think about it is:

Network partition occurs
          |
          v
   ┌──────┴──────┐
   |             |
   v             v
Preserve       Preserve
Consistency    Availability
   |             |
   v             v
Reject         Serve
requests       requests

The system's behavior during the partition determines the tradeoff.


Why Is Partition Tolerance Usually Considered Mandatory?

If you're building a distributed system across machines or data centers, network failures are inevitable.

Servers crash.

Routers fail.

Connections drop.

Packets disappear.

Data centers lose connectivity.

You can't simply tell the network:

"Please never partition."

So real distributed systems generally have to assume that partitions can happen.

That leaves the practical question:

When communication breaks, do we stop serving some requests or continue serving them with potentially stale data?

That's the heart of CAP.


A Banking Example

Imagine a banking system.

You have:

Bank DC 1
    |
    X
    |
Bank DC 2

The network connection between the two data centers breaks.

A customer has:

$1,000

in their account.

Now they try to withdraw $900 from DC 1.

At almost the same time, another transaction tries to withdraw $900 through DC 2.

If both sides continue operating independently:

DC 1 thinks:
 
$1,000 → $100
 
 
DC 2 thinks:
 
$1,000 → $100

Together...

The customer just spent $1,800.

But they only had $1,000.

This is exactly the kind of scenario where sacrificing availability can be safer than accepting inconsistent state.


Now Consider Social Media

Suppose you like a post.

The system has thousands of replicas around the world.

Your request reaches one server:

likes = 10,000

Your like gets accepted:

likes = 10,001

Another user in another region might temporarily see:

10,000

instead of:

10,001

Is that catastrophic?

Probably not.

Eventually:

10,000
   ↓
10,001

The system becomes consistent again.

That's eventual consistency.

The application chooses availability and accepts temporary inconsistency because the consequences are small.


Strong Consistency vs Eventual Consistency

This distinction is important.

Strong consistency

After a successful write:

Write X
 
↓
 
Every subsequent read
sees X

Eventual consistency

After a successful write:

Write X
 
↓
 
Some replicas see X
 
Some replicas don't
 
↓
 
Replication continues
 
↓
 
Eventually everyone sees X

Neither is universally "better."

It depends on the application.


CAP Is Not PACELC

Here's where distributed systems get even more interesting.

CAP only describes what happens during a partition.

But what about when there isn't a partition?

A different model called PACELC asks another question:

If there is a Partition, do you choose Availability or Consistency? Else, do you choose Latency or Consistency?

In simplified form:

        Partition?
        /        \
      Yes         No
       |           |
     A or C      L or C

Because even when the network is healthy, distributed systems still face tradeoffs.

For example:

Do I wait for another replica to acknowledge the write?

That may improve consistency.

But it also increases latency.

So distributed systems are full of tradeoffs even outside the CAP scenario.


The Most Important Mental Model

Don't memorize:

CAP = Pick 2

Remember this instead:

Normal operation:
 
C + A + P can coexist
 
        ↓
 
Network partition occurs
 
        ↓
 
C  <───────>  A
 
You cannot guarantee both

That's a much better mental model.


One More Important Detail

CAP's definition of Consistency is not the same thing as database consistency in ACID.

This causes a lot of confusion.

CAP consistency is essentially about linearizability—whether operations appear to occur atomically in a single, real-time-consistent order.

ACID's "C" stands for something different: maintaining application-defined integrity constraints.

Same letter.

Different concept.


Why Engineers Should Actually Care

CAP isn't just an interview question.

It forces you to ask a practical question when designing distributed systems:

What happens when our replicas can't communicate?

For every important operation, ask:

Network partition occurs.
 
What do we do?
 
        ↓
 
Reject request?
 
        OR
 
Serve potentially stale data?

For a payment authorization?

You probably want very strong guarantees.

For a social-media like count?

Temporary inconsistency might be completely acceptable.

For a shopping cart?

You might choose a more nuanced strategy depending on the operation.

The correct answer depends on the business requirement.


Final Thoughts

CAP Theorem isn't:

"Pick any two."

It's a statement about an unavoidable tradeoff when a network partition prevents distributed nodes from communicating.

When everything is healthy, your system can provide strong consistency and high availability.

But when the network breaks...

You have to decide what matters more.

             Network Partition
                    |
             ┌──────┴──────┐
             ↓             ↓
        Consistency    Availability
             |             |
        "Don't guess"   "Keep serving"
             |             |
             ↓             ↓
          Fail some     Accept possible
          requests      stale data

The best distributed systems aren't the ones that blindly choose "CP" or "AP."

They're the ones that understand which guarantees each operation actually needs.

So the next time someone says:

"CAP means choose any two."

Ask them one question:

"During what condition?"

That's where the real theorem begins.