All notes

When Everything Is Green, But Your App Is Still Down

A successful deployment doesn't always mean a healthy application. Here's how senior engineers debug incidents when every dashboard says everything is fine.

4 min read

Imagine you just deployed your application.

CI/CD passed.

The deployment completed successfully.

The server is running.

Your Kubernetes pods are healthy.

The health endpoint returns 200 OK.

Everything looks green.

Yet users keep reporting:

"The website isn't working."

Where do you even start?

The Biggest Mistake

Many engineers immediately SSH into the server.

Or worse...

They redeploy.

A senior engineer starts with one simple question.

"Can I reproduce the problem?"

If users can't access the application but you can, the problem might not even be inside your backend.

Step 1: Is It Really Down?

The first thing to verify is whether the issue affects everyone.

Can employees access it?
 
      Yes
 
Can customers access it?
 
      No

If only external users are affected, the application itself may be perfectly healthy.

Step 2: Health Checks Don't Mean Your App Works

A health endpoint often checks something simple.

GET /health
 
200 OK

Behind the scenes, it might only verify:

return "OK";

or

database.ping();

That tells you the application is alive.

It doesn't prove your login page, APIs, authentication, or frontend are functioning correctly.

A healthy server does not guarantee a healthy product.

Step 3: Follow The User's Request

Instead of looking at logs randomly, trace one failing request.

Browser
 
    |
 
CDN
 
    |
 
Load Balancer
 
    |
 
API Gateway
 
    |
 
Backend
 
    |
 
Database

Somewhere along this path, the request is failing.

The goal is to find exactly where.

Step 4: Check The Browser

Many production incidents aren't backend problems at all.

Open DevTools.

Look for:

404
 
401
 
403
 
500
 
CORS Errors
 
Mixed Content
 
Failed Network Requests

A React application may load successfully while every API request silently fails.

Step 5: CDN And Cache

Sometimes the deployment succeeded...

But users are still receiving yesterday's frontend bundle.

Browser
 
↓
 
CDN Cache
 
↓
 
Old JavaScript

The browser requests an API that no longer exists.

Frontend
 
GET /api/v1/users
 
Backend
 
/api/v2/users

Everything works individually.

Together, they break.

Step 6: DNS And Load Balancer

Your application might be running perfectly.

Traffic simply isn't reaching it.

Possible causes include:

The backend never even receives the request.

Step 7: Check Dependencies

Modern applications rarely work alone.

Frontend
 
↓
 
API
 
↓
 
Redis
 
↓
 
Database
 
↓
 
Authentication Service

Suppose your authentication provider is unavailable.

The API still returns:

GET /health
 
200 OK

Yet every user receives:

Login Failed

The application isn't technically down.

One dependency is.

Step 8: Verify The Deployment

Sometimes only half the deployment succeeded.

Imagine two application servers.

Load Balancer
 
     |
 
-------------
 
Server A
 
Version 2.0 ✅
 
-------------
 
Server B
 
Version 1.0 ❌

Half the users hit the new version.

The other half hit the old version.

Only some users experience failures.

These issues are notoriously difficult to spot.

Step 9: Use Distributed Tracing

Instead of reading thousands of log lines, trace a single request.

Client
 
↓
 
Gateway
 
↓
 
User Service
 
↓
 
Order Service
 
↓
 
Database

Tracing immediately reveals where latency or failures occur.

Without tracing, you're mostly guessing.

How Senior Engineers Think

Notice something interesting.

They don't start by looking at code.

They start by narrowing the problem space.

Each answer eliminates an entire category of possible failures.

Final Thoughts

A deployment succeeding only means new code reached production.

It says nothing about whether users can actually use your application.

The best engineers don't panic when every dashboard is green.

They follow the request, eliminate possibilities one by one, and let evidence guide the investigation instead of assumptions.

Debugging production systems isn't about finding bugs quickly.

It's about asking the right questions in the right order.