All notes

The Performance Bug That Only Happens 2% of the Time

Most APIs are optimized for average latency. Senior engineers optimize for tail latency, where a small percentage of requests become unexpectedly slow.

4 min read

Imagine you own an API.

49 requests complete in 50ms.

But every once in a while, one request suddenly takes 800ms.

No exceptions.

No errors.

No increased CPU usage.

No database failures.

Just one request randomly becoming much slower than the rest.

Most engineers immediately start looking for the average response time.

The problem is that averages hide the worst experiences.

Request Times
 
50ms
48ms
52ms
49ms
51ms
780ms šŸ‘ˆ
47ms
53ms

The average still looks healthy.

But one unlucky user waited over 15 times longer.

This phenomenon is called Tail Latency.

What Is Tail Latency?

Tail latency measures the slowest requests rather than the average.

Instead of asking:

"What's our average response time?"

we ask:

"How long do the slowest 1% of requests take?"

This is why production dashboards often display:

P50   45ms
 
P95   120ms
 
P99   780ms

If your P99 is high, some users are having a poor experience even though the average looks excellent.

Where Does Tail Latency Come From?

Many engineers immediately blame the database.

Sometimes that's correct.

Often it isn't.

A single slow request could be caused by:

The challenge is that these issues don't happen consistently.

They only affect a small percentage of requests.

The Wrong Way To Debug

Many people start with application logs.

System.out.println("API Called");

Unfortunately, logs rarely explain why one request suddenly became slow.

They only tell you that it happened.

The Better Approach: Distributed Tracing

Every request receives a unique Trace ID.

Client
 
   |
 
API Gateway
 
   |
 
User Service
 
   |
 
Payment Service
 
   |
 
Database

Instead of guessing, you can see exactly where time was spent.

Request
 
API Gateway         5ms
 
User Service        12ms
 
Payment Service    620ms šŸ‘ˆ
 
Database            18ms

Now the bottleneck is obvious.

Without tracing, engineers often waste hours looking in the wrong place.

Thread Pool Starvation

Another common cause is thread exhaustion.

Imagine your application has only ten worker threads.

ExecutorService executor =
        Executors.newFixedThreadPool(10);

If all ten threads are waiting on slow database calls, the eleventh request must wait.

Thread Pool
 
[Busy]
[Busy]
[Busy]
[Busy]
...
[Busy]
 
Incoming Request
 
↓
 
Waiting...

The API isn't slow because the code is inefficient.

It's slow because there are no available worker threads.

Cache Misses

Caching creates another interesting scenario.

Request 1
 
Redis Cache
 
MISS āŒ
 
↓
 
Database
 
↓
 
Cache Updated
 
------------------------
 
Request 2
 
Redis Cache
 
HIT āœ…

Most users experience fast responses.

Occasionally, a cache miss forces the application to query the database, increasing latency for that request.

Lock Contention

Concurrency can also introduce tail latency.

Suppose multiple threads attempt to update the same record.

synchronized(lock){
 
    updateAccount();
}

Only one thread enters the critical section.

Everyone else waits.

As traffic increases, waiting time grows.

Garbage Collection

Even if your application code is perfect, the JVM may temporarily pause execution.

During a major garbage collection cycle:

Application Running
 
↓
 
GC Starts
 
↓
 
Application Pauses
 
↓
 
GC Finishes
 
↓
 
Application Resumes

Most pauses last only milliseconds, but large heaps can occasionally produce noticeable latency spikes.

How Senior Engineers Debug This

Instead of asking:

"Why is the API slow?"

They ask:

"Why are only some requests slow?"

Their investigation usually includes:

The goal isn't simply finding a slow API.

It's identifying which component contributes to the tail.

Final Thoughts

Most applications are optimized for average performance.

Large-scale systems are optimized for worst-case performance.

Users don't remember the 49 fast requests.

They remember the one request that took forever.

That's why senior engineers spend far more time analyzing tail latency than average latency.

In distributed systems, the slowest request is often the one that teaches you the most.