All notes

Why Counting Active Sessions Isn't Enough: How Netflix Enforces Device Limits at Scale

Limiting users to two devices sounds simple until millions of devices hit Play simultaneously. Here's how large streaming platforms solve race conditions, session tracking, and distributed consistency.

6 min read

Enforcing a 2-device streaming limit sounds like one of the easiest backend problems.

Until you have 300 million users.

Imagine this.

A Premium plan allows only 2 simultaneous streams.

Currently:

TV ✅
 
Laptop ✅

Now the user opens Netflix on:

Phone
 
Tablet

Both devices press Play within milliseconds.

How do you make sure only one of them starts streaming?

Most candidates answer something like this.

SELECT COUNT(*)
FROM active_sessions
WHERE user_id = 101;
 
IF count < 2
 
INSERT INTO active_sessions(...)
 
ELSE
 
Reject

Looks correct.

Unfortunately...

It breaks almost immediately in production.

The Problem Starts With Race Conditions

Suppose two requests arrive simultaneously.

Phone
 
↓
 
Server A
 
------------------------
 
Tablet
 
↓
 
Server B

Server A executes:

COUNT(active_sessions)
 
↓
 
2

Before it inserts...

Server B also executes:

COUNT(active_sessions)
 
↓
 
2

Both believe there's one slot available.

Both insert.

Now the user has:

TV
 
Laptop
 
Phone
 
Tablet

Four active streams.

The limit was two.

This isn't a rare edge case.

At Netflix's scale, thousands of these collisions happen every second.

The problem is that SELECT followed by INSERT is not atomic.

Atomic Check And Insert

The critical operation isn't counting.

It's ensuring that:

Read
 
↓
 
Validate
 
↓
 
Insert

happens as one indivisible operation.

One common solution is Redis.

Instead of splitting the work into multiple requests, Redis executes everything inside a single Lua Script.

Read Current Sessions
 
↓
 
Is Limit Reached?
 
↓
 
Yes → Reject
 
↓
 
No → Add Session

Since Redis executes Lua scripts atomically, no other request can modify the session count midway.

Only one device wins.

The other sees the updated state.

Why Redis?

Redis is extremely fast.

Operations complete in microseconds.

More importantly, atomic commands eliminate race conditions without requiring heavyweight database locks.

This keeps the hot path extremely efficient.

But Rejecting Users Isn't Great UX

Suppose the user forgot to close Netflix on an old Smart TV yesterday.

Should the new playback fail?

Probably not.

A better user experience is:

Current Sessions
 
TV (Idle)
 
Laptop (Watching)
 
----------------------
 
Phone Starts Playing

Instead of rejecting the Phone...

Silently remove the idle TV session.

Phone ✅
 
Laptop ✅
 
TV ❌

The user never notices.

Only when every active device is genuinely streaming should the new request be rejected.

Business logic matters just as much as technical correctness.

What Counts As "Active"?

Another subtle problem.

Someone pauses a movie.

Leaves for lunch.

Never closes Netflix.

Should that session count forever?

Obviously not.

Large streaming platforms usually rely on heartbeats.

Every device periodically sends a lightweight request.

Every 30 Seconds
 
↓
 
POST /heartbeat

Each heartbeat refreshes the session's expiration time.

Redis
 
Session
 
TTL = 30 Seconds

As long as heartbeats continue, the session remains active.

If they stop...

Redis automatically removes the session after the TTL expires.

No cleanup jobs.

No scheduled database tasks.

No cron jobs.

Everything expires naturally.

Why Not Store Sessions In MySQL?

Imagine millions of devices sending heartbeats every 30 seconds.

That would generate an enormous number of database writes.

Instead, Redis handles the hot path.

The database is updated asynchronously only for:

This separation keeps the primary database free for business-critical operations.

What Happens If Redis Goes Down?

Interviewers love asking this.

Suppose Redis becomes unavailable for ten seconds.

Should Netflix block every paying customer?

Probably not.

Most production systems would fail open.

Redis Down
 
↓
 
Temporarily Allow Playback

A handful of extra streams for a few seconds is far less damaging than preventing millions of legitimate customers from watching.

Availability is more important than perfect enforcement.

Since Redis typically runs with replication and automatic failover, outages are usually measured in seconds.

What About Users Who Kill The App?

Another edge case.

Suppose someone force closes Netflix.

No logout.

No heartbeat.

No disconnect event.

Without TTLs, that session would remain active forever.

Heartbeats solve this naturally.

Heartbeat Stops
 
↓
 
TTL Expires
 
↓
 
Session Removed

No manual cleanup required.

A Typical Flow

Device Presses Play
 
↓
 
Redis Lua Script
 
↓
 
Check Active Sessions
 
↓
 
Under Limit?
 
      │
 
 Yes  │  No
 
      ▼
 
Create Session
 
      │
 
Heartbeat Every 30 Seconds
 
      │
 
User Stops Watching
 
      │
 
Heartbeat Stops
 
      │
 
TTL Expires
 
      │
 
Session Removed

Follow-Up Questions Interviewers Love

Why Redis instead of MySQL?

Because session enforcement is on the critical request path and requires atomic operations with extremely low latency. Redis provides both.


Why not use database transactions?

Database locks don't scale well for millions of concurrent session checks. They're slower and create contention.


Why use Lua Scripts?

Multiple Redis commands executed separately can still introduce race conditions.

Lua Scripts execute atomically.


Why Heartbeats instead of Logout?

Users frequently lose internet, force close applications, or switch devices.

Heartbeats ensure inactive sessions disappear automatically without relying on the client.


Why Fail Open instead of Fail Closed?

Blocking every legitimate customer because one dependency is temporarily unavailable creates a far worse experience than allowing a few extra sessions for a few seconds.

Understanding this tradeoff is what production engineering is all about.

Lessons Beyond Netflix

This interview isn't really about streaming.

It's testing whether you recognize:

Many real-world systems share the same pattern.

Different products.

The same distributed systems problems.

Final Thoughts

At first glance, limiting a user to two devices feels like a simple COUNT(*) query.

At production scale, it's a distributed systems problem involving concurrency, atomicity, caching, session lifecycle management, and business tradeoffs.

The best engineers don't immediately reach for Redis, transactions, or locks.

They first identify what absolutely must remain correct, and then apply the simplest solution that guarantees that correctness.

That's exactly what interviewers are looking for.