ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 51

  1. Blog

How to Read Server Monitoring Graphs

A framework for turning dashboard noise into a diagnosis: symptoms first, causes second, and why averages lie about response time.

Aditya Challa·22 September 2026·7 min read

Summary
On this page
  1. The four signals everything reduces to
  2. Traffic: context before symptom
  3. Latency: why the average lies
  4. Errors: which kind, and how fast they fail
  5. Narrowing down the cause: CPU as an unreliable narrator

Most monitoring guides cover how to install and build dashboards. Almost none cover how to actually read them once they're built. That gap is the subject of a guide from kciter.so, which lays out a structured way to move from "the dashboard looks weird" to "here's the cause."

The four signals everything reduces to

No matter how many panels a dashboard has, every metric on it falls into one of four categories — what Google's SRE organization calls the golden signals, per the source:

  • Traffic: how much is coming in
  • Latency: how long does it take
  • Errors: how often requests fail
  • Saturation: how full the server's resources are

The guide adds a second axis on top of this: metrics split into symptoms and causes. Latency and error rate are symptoms — they're what users actually experience, and when they rise, something is genuinely wrong. Resource metrics like CPU, memory, and thread pools are causes. The source is explicit that a resource number on its own doesn't tell you anything: "CPU at 90% is not an immediate emergency if responses are fast, and CPU at 20% is a problem if responses are slow." A sustained 90% is worth watching because it means no headroom is left, but by itself it isn't a reason to page anyone.

The practical rule that follows: identify the symptom first, then use the cause-side metrics to narrow down why. Alerts, the source argues, should be attached to symptoms, not resource thresholds — a brief "CPU over 80%" spike wakes someone for nothing, but an "error rate exceeded" alert means someone is actually having a bad time.

Traffic: context before symptom

Strictly speaking traffic isn't a symptom, it's context — but you need it to interpret latency and errors correctly. The first skill, according to the guide, is memorizing your service's normal shape. Without that baseline, a metric that's technically inside its "normal range" but at half its usual level won't register as an anomaly.

Once you know the shape, a few anomaly patterns cover most cases:

  • Traffic dropping vertically doesn't mean the server went idle — it usually means requests are failing to arrive upstream, at a load balancer, DNS, or gateway. If every server-side graph looks calm during an incident, that calm is itself the anomaly.
  • Traffic shooting up vertically points to a viral surge, crawler traffic, or an attack.
  • Regular spikes at a fixed time (early morning, for instance) are almost always batch jobs or cron — if the pattern is periodic, suspect an internal job first.

Traffic also functions as a denominator: 500 errors mean something completely different at a million requests per minute versus a thousand. Any raw count — errors, slow queries, whatever — needs to be read as a rate against traffic, not as an absolute number.

Latency: why the average lies

The source poses a scenario worth sitting with: average response time has been flat at 100ms for days, but users keep complaining the app is slow. The dashboard says healthy, the users say problem. The average is the one lying, because it erases the shape of the underlying distribution. A server where every request takes ~100ms and a server where most take 30ms but some take 900ms can produce the identical average.

That's why the guide insists on reading latency in percentiles — P50, P95, P99 — rather than averages. P99 in particular gets called "tail latency," and the source gives two reasons not to write it off as a rounding error:

  1. At any real scale, 1% is a lot of requests. The source's own example:
# At 1,000 requests/sec, P99 = 1% missing the fast path
requests_per_second = 1000
p99_fraction = 0.01
users_hitting_p99_per_second = requests_per_second * p99_fraction  # 10
users_hitting_p99_per_day = users_hitting_p99_per_second * 86400   # 860,000
  1. Modern pages fire dozens of API calls to render one screen, and the odds of hitting the tail compound across calls:
# Probability a page avoids P99 across N sequential calls
p_avoid_single_call = 0.99
n_calls = 40
p_avoid_all_calls = p_avoid_single_call ** n_calls  # ≈ 0.67

Even if a single call avoids the tail 99% of the time, forty calls together only avoid it about 67% of the time — meaning roughly one user in three hits tail latency somewhere on the screen. The source also notes that tail-latency users aren't random: heavy, long-time users generate heavier queries and land in P99 more often, so "P99 is quite possibly the response time your best customers are getting." The recommendation: P50 as the representative user-experience number, P95 or P99 as the basis for alerts and performance targets, since early signs of an incident tend to show up in P99 first.

Errors: which kind, and how fast they fail

When an error graph jumps, the guide says to ask "which errors" before "how many." A 5xx is always the server's problem. A 4xx is technically a client mistake by spec, but a surge of 400s or 401s right after a deploy is more likely a broken API contract — a new client version talking to an old server, or vice versa — than a sudden wave of bad clients.

The second thing to check is how fast the failures happen. Timeouts fail slowly: they hold a thread and a connection for seconds before giving up, hitting error rate and saturation at once. Connection refusals or code bugs like null references fail immediately, returning a 5xx within milliseconds and freeing resources instead of holding them. Reading error rate against P99 together tells you which kind you're facing — both rising together suggests something slowing down and dying; error rate spiking while P99 stays flat suggests something failing instantly. The source flags a specific trap here: requests that fail instantly can leave the latency distribution entirely, so the worse an incident gets, the better P99 can look. An incident where error rate is worst but response times look better than usual is a sign of survivorship bias, not recovery.

Narrowing down the cause: CPU as an unreliable narrator

Once the symptom is identified, the guide moves to cause-side metrics — starting with the warning that a resource metric read alone will lie to you. CPU utilization is usually the first panel on any dashboard, and it's described as "surprisingly uninformative" on its own. The fix: read it overlaid with response time, which produces three distinct diagnoses from the same CPU shape:

  • Healthy: CPU rises and falls with the traffic curve, response time stays low and stable.
  • CPU idle at 20–30% while response time explodes: counterintuitively, this usually means trouble. Idle CPU with slow responses means threads are waiting, not computing — a slow DB, a held lock, an exhausted connection pool, or a stalled external API. The fix is to look past CPU toward I/O and pool metrics.
  • CPU pinned flat at 100%: either traffic has genuinely exceeded capacity (if traffic rose with CPU — scale out) or code is spinning in a loop (if traffic is normal and only CPU spiked — go look at the code).

The source also flags a container-specific wrinkle: in Kubernetes, a CPU limit means the container can only use its quota per period (100ms by default), and once that's used up it's forcibly paused until the next period — a behavior called throttling. If traffic and code are unchanged but the latency tail has grown, this setting is worth checking before anything else.

The source material stops short of covering the guide's remaining sections on bottlenecks, backpressure, caches, and timeouts — those weren't in the excerpt available here. Everything above summarizes the linked article as written.

Sources

  1. 01
    How to Read Server Monitoring Graphs

    For a server developer, monitoring is something like a lifeline. It has to be, because a server is a time bomb that can go off at any moment. With monitoring in place, you can narrow down the cause quickly when an incident happens, and sometimes catch a problem before it becomes one. Unfortunately, while there are plenty of guides on installing and building monitoring dashboards, there aren’t many on how to read those dashboards and diagnose problems with them. So a lot of developers feel a vag

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 42

    Chrome DevTools AI Assistance (Gemini): Enable + Prompt Guide 2026

    1 Oct 2026

  • Plate 46

    Remix 3 Ships Oct 2, 2026: What Frontend Teams Need to Know

    1 Oct 2026

  • Plate 87

    Pennsylvania Measles Outbreak FAQ (October 2026)

    Factual FAQ on Pennsylvania’s 2026 measles outbreak: case counts, deaths, MMR basics, and where to find official CDC and state health updates.

    1 Oct 2026

On this page

  1. The four signals everything reduces to
  2. Traffic: context before symptom
  3. Latency: why the average lies
  4. Errors: which kind, and how fast they fail
  5. Narrowing down the cause: CPU as an unreliable narrator
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove