DEV Community

Cover image for Your App Being “Online” Doesn’t Mean It’s Healthy
The Saint
The Saint

Posted on

Your App Being “Online” Doesn’t Mean It’s Healthy

One of the easiest traps in production is assuming:

“The website is responding, so everything is fine.”

It isn't.

An application can technically be online while quietly falling apart underneath.

Your API might still respond while CPU usage sits at 95%.

Your server might still accept traffic while disk space has 300 MB left.

Your database might be reachable while requests are taking five seconds instead of 200 milliseconds.

Your process might still be running while one important background worker died twenty minutes ago.

And sometimes the first person to tell you something is wrong is a user.

That's usually too late.

Uptime is only one signal

The simplest form of monitoring asks one question:

Is the application reachable?

You send a request every few seconds or minutes.

If it returns successfully:

everything looks green.

If it stops responding:

send an alert.

That is useful.

But it only detects one class of failure.

Production systems often degrade long before they completely disappear.

Imagine this:

Technically, the server is alive.

Operationally, you should probably be worried.

Monitoring the application

Application monitoring looks at what is happening inside the software.

Things like:

application crashes
unhandled exceptions
failed requests
response time
API errors
runtime failures
unhealthy processes
failed jobs

Consider an API that normally responds in:

and gradually moves to:

Nothing has technically crashed.

But something has changed.

Maybe a database query became slower.

Maybe traffic increased.

Maybe a dependency is struggling.

Maybe memory pressure is causing garbage collection problems.

Monitoring gives you a chance to investigate before the application becomes unavailable.

Then there is the machine underneath it

This is the part I think developers sometimes forget.

Your application isn't floating in the cloud.

Eventually, it is running on a machine somewhere.

That machine has:

And any one of those can become the actual reason your application fails.

For example, your Node.js API might be perfectly written.

but It doesn't matter if the VPS disk is full.

Your application can't write new files.

Your database may not be able to expand.

Logs stop being written.

Eventually, things start breaking in strange ways.

The application is the symptom.

The server is the cause.

CPU spikes matter

Suppose your server normally runs around:

Then suddenly:

for several minutes.

That might mean:

unexpected traffic
runaway process
expensive database operation
infinite loop
background job
compromised server
badly optimized code

You don't necessarily want an alert every time CPU touches 90% for three seconds.

That would become noise.

But sustained abnormal usage?

You probably want to know.

Something like:

CPU usage has remained above 90% for 5 minutes.

That's actionable.

Disk space is boring until it destroys your weekend

Disk monitoring might be one of the least exciting things in software.

Until this happens:

Databases need disk.

Logs need disk.

Uploads need disk.

Temporary files need disk.

Docker images need disk.

Backups may need disk.

A slowly growing log directory can quietly consume storage for weeks before anything obviously fails.

So instead of discovering:

after the damage starts, monitoring should warn you earlier:

The alert isn't the solution.

It gives you time to create one.

Network problems are different from downtime

A server can be reachable and still have terrible connectivity.

Maybe:

Your users experience the application as slow.

Your uptime monitor says:

All systems operational.

Both statements can technically be true.

This is why latency belongs beside uptime.

You want to understand:

Is the server reachable?
How quickly is it responding?
Has network latency changed?
Are packets being dropped?
Is one region behaving differently?

Availability and performance are related, but they are not the same thing.

Monitoring without alerts is just a dashboard

There is another mistake I've seen:

Beautiful monitoring dashboards.

Charts everywhere.

CPU graphs.

Memory graphs.

Latency graphs.

Error graphs.

But nobody is staring at those dashboards at 3:17 AM.

A monitoring system needs a way to say:

Something changed. You should look at this.

That could be:

depending on how the team works.

The important part is delivering the right alert to somebody capable of acting on it.

For example:

Production API Down

api.example.com has failed 3 consecutive health checks.

Last successful check: 02:41 UTC

Or:

Server CPU Alert

CPU usage on production-01 has remained above 90% for 5 minutes.

That is much more useful than discovering the problem because somebody posted:

“Is your site down?”

on Twitter.

But alerts can also become useless

The opposite problem is alert fatigue.

If your system tells you:

every thirty seconds, you eventually stop caring.

Then the important notification gets buried with everything else.

Good monitoring needs thresholds.

And often, duration.

Instead of:

CPU crossed 80%.

think:

CPU remained above 85% for five minutes.

Instead of:

One request failed.

think:

Error rate exceeded 5% over the last five minutes.

The goal isn't maximum notifications.

The goal is useful awareness.

Monitoring should exist at multiple layers

I've started thinking about production monitoring as two connected systems.

Layer 1 — Application

Watch:

Layer 2 — Infrastructure

Watch:

Because sometimes the application is broken.

Sometimes the infrastructure is broken.

And sometimes one is telling you something about the other.

Seeing both makes diagnosis much easier.

Monitoring shouldn't begin after something breaks

This is probably the most important part.

Teams often add monitoring after their first serious incident.

Something goes down.

Nobody notices for forty minutes.

Users complain.

The team fixes it.

Then someone says:

“We should probably add monitoring.”

Monitoring shouldn't be the postmortem action item.

It should be part of getting something into production.

I've increasingly started thinking about the backend lifecycle as:

Build → Verify → Deploy → Operate → Own → Monitor

Building answers:

Can we create it?

Verification asks:

Does it behave correctly?

Deployment asks:

Can we get it into production?

Operations asks:

Can we manage it there?

Ownership asks:

Do we retain control of what we're running?

And monitoring asks:

Will we know when something changes?

That last question matters more once nobody is sitting there watching the terminal.

This is something we're working on ourselves

We've been thinking about this while building CrescoDB.

CrescoDB already manages the production server layer for projects deployed through it, so we can see things like server availability and runtime health.

We're now expanding that into a more complete monitoring layer.

Not an attempt to recreate every observability feature from platforms built entirely around monitoring.

The goal is simpler:

Give smaller engineering teams the information they actually need to know when something is going wrong.

At the application level:

And underneath that:

Then push useful alerts through the channels teams are already watching:

Telegram, Slack and email.

The interesting part for us isn't adding another dashboard.

It's connecting application health with the infrastructure running it.

Because if your API goes down at the same moment the server's memory hits 98%, those two pieces of information probably belong beside each other.

Production is really about feedback

When you're developing locally, feedback is immediate.

Something breaks.

Your terminal tells you.

You fix it.

Production creates distance between the developer and the software.

Monitoring closes some of that distance.

It tells you:

Something changed.

before the user has to.

And that might be the simplest definition of good monitoring:

Knowing your software is unhealthy before your customers do.

What production issue have you had that proper monitoring would have caught earlier?

Top comments (0)