Plate 44
OpenTelemetry Collector + Prometheus + Grafana: A First Stack That Actually Diagnoses
A first OTel Collector, Prometheus, and Grafana lab: OTLP in, Prometheus scrape out, and a break-one-signal drill with measured PromQL.
Aditya Challa10 min read
On this page
- Intro — what this post promises
- What each piece owns
- Topology
- Collector config — the pieces that mattered
- Three ways metrics reach Prometheus — we ran Path A
- Grafana — the minimum dashboard that diagnosed
- What the demo emitted
- The break-one-signal drill
- Honest advantages and disadvantages
- Footprint on this box (not a 30-minute soak)
- FAQ
- More reading
Intro — what this post promises
Most “OTel + Prometheus + Grafana” tutorials stop when the dashboard turns green. The bar here is a stack that diagnoses: when the app dies, when the SDK dials the wrong Collector port, when the Collector itself stops — which panel moves, and which one stays calm on purpose?
This is the opinionated minimal lab for that question:
- Why a Collector belongs in the path.
- The topology (app → OTLP → Collector → Prometheus → Grafana).
- One metrics path, run for real: Collector Prometheus exporter, scraped.
- A break-one-signal drill with timestamps from 29 Sep 2026 (IST).
- Where percentiles and golden signals live — other posts, not this one.
Golden-signal literacy stays on How to read server monitoring graphs. Histogram and quantile rules stay on Why your average latency graph is lying (CMS auto-slug until the slug editor ships). This post only hosts the panels those playbooks query.
Related links:
Lab honesty: Docker was not installed on the lab box (docker not found). The same pinned versions ran as host binaries, not containers: OpenTelemetry Collector contrib 0.111.0, Prometheus 2.54.1, Grafana 11.2.0, and a tiny Go 1.24.4 demo using the OpenTelemetry Go SDK v1.32.0 (otlpmetricgrpc v1.32.0). Path used: A — prometheus exporter + scrape. otelcol-contrib validate exited 0. This was a few minutes of load on a shared Linux box, not a 30-minute idle soak and not a production cluster.
A practical first OpenTelemetry stack is: app → OTLP → Collector → Prometheus → Grafana. The Collector receives OTLP, batches, and exposes a Prometheus text endpoint. Grafana queries Prometheus. Prove the stack by breaking one signal and watching the right pillar fail.
What each piece owns
| Component | Job in this lab | Where to start |
|---|---|---|
| App / SDK | Emit metrics via OTLP/gRPC | Language OTel SDK docs |
| OpenTelemetry Collector | Receive, batch, export | Collector docs |
| Prometheus | Store and query metrics (PromQL) | Prometheus scrape config |
| Grafana | Dashboards on top of Prometheus | Prometheus data source |
OpenTelemetry’s Collector overview recommends a Collector next to services in real deployments: apps offload quickly; the Collector handles retries, batching, and filtering. Direct SDK → backend is fine for a tiny experiment. This lab still uses a Collector so the shape matches production.
Related links:
Topology
Traces and logs are out of scope. Ship metrics first so the percentiles playbook has a real histogram to query.
Related links:
Listeners in this lab were bound to 127.0.0.1, not 0.0.0.0. Collector docs warn that an open OTLP port is a denial-of-service surface. On a laptop that is the right default; on a VPS, firewall the port or keep it on a private interface.
Collector config — the pieces that mattered
Collector config is always the same spine (configuration guide): receivers → processors → exporters, then service.pipelines that actually enable them. A component that is not listed in a pipeline does nothing.
Related links:
This is the file that validated and ran (contrib 0.111.0):
namespace: lab is why the demo’s demo.requests counter showed up in Prometheus as lab_demo_requests_total. Resource attributes (service.name, deployment.environment) were copied onto the series because resource_to_telemetry_conversion was on.
Three ways metrics reach Prometheus — we ran Path A
Interoperability is documented from both sides (OTel “Prom and OTel”, OTLP metrics export to Prometheus):
Related links:
| Path | How it works | This lab? |
|---|---|---|
| A. prometheus exporter + scrape | Collector serves Prom text; Prometheus scrapes :8889 | Yes — this is what ran |
| B. prometheusremotewrite | Collector pushes to Prom’s remote-write receiver | Not run |
| C. OTLP → Prometheus | Push OTLP HTTP to Prom’s native OTLP metrics endpoint | Not run |
Path A is the right first lab because the mental model is still “Prometheus scrapes a target.” Path C is the more OTel-native production story on modern Prometheus; do not implement all three in one afternoon.
Prometheus config that scraped the Collector (interval 5s):
Working PromQL at 23:10 IST (steady demo load, 2-minute rate window):
| Query | Result |
|---|---|
up{job="otel-collector"} | 1 |
sum(rate(lab_demo_requests_total[1m])) | 2.68 req/s |
sum(rate(lab_demo_errors_total[1m])) | 0.105 req/s |
histogram_quantile(0.50, sum by (le) (rate(lab_demo_latency_seconds_bucket[2m]))) | 28.6 ms |
| same, φ=0.95 | 48.1 ms |
| same, φ=0.99 | 49.9 ms |
lab_demo_requests_total | 169 (HTTP 200) and 7 (HTTP 500) at that instant |
The cumulative counter later froze at 473 when the demo stopped (see the drill). Series carried service_name="shoppercove-demo", deployment_environment="lab", and http_route="/work".
If you would rather ship containers, pin the same versions. This Compose file was not executed here (no Docker). It matches the binaries that were:
Grafana — the minimum dashboard that diagnosed
Prometheus data source: http://127.0.0.1:9090 (inside Compose that URL is http://prometheus:9090). Grafana 11.2.0 provisioned five panels:
- Request rate —
sum(rate(lab_demo_requests_total[1m])) - Error rate —
sum(rate(lab_demo_errors_total[1m])) - Latency quantiles —
histogram_quantileonlab_demo_latency_seconds_bucketfor p50 / p95 / p99. Rules for when those numbers lie are in the percentiles post, not here. - Scrape health —
up{job=~"otel-collector|prometheus"} - Raw counter —
lab_demo_requests_total(the flatline is easier to trust than a rate window)
Related links:
From process start (23:08 IST) to a Grafana panel that showed request rate (23:11 IST) was about two minutes. The Collector’s :8889/metrics already listed lab_demo_* about 45 seconds after the demo started.
What the demo emitted
One Go process, not a microservice estate. OpenTelemetry Go SDK v1.32.0, exporter otlpmetricgrpc v1.32.0, gRPC v1.67.1, runtime Go 1.24.4.
| Signal | OTel name | Prometheus name (namespace lab) |
|---|---|---|
| Request counter | demo.requests | lab_demo_requests_total |
| Error counter | demo.errors | lab_demo_errors_total |
| Latency histogram | demo.latency (unit seconds, explicit buckets 5 ms … 10 s) | lab_demo_latency_seconds_bucket |
| Resource | service.name=shoppercove-demo, deployment.environment=lab | labels on the series |
Export interval was 2 seconds to 127.0.0.1:4317. Proof the series existed: the PromQL table above, and GET http://127.0.0.1:8889/metrics returning lab_demo_requests_total{http_status_code="200",...}.
The break-one-signal drill
All times 29 Sep 2026, Asia/Calcutta. Scrape interval 5 seconds.
| Break | When (IST) | What moved | What stayed | Pass? |
|---|---|---|---|---|
| Stop the demo process | Killed 23:12:18. At 23:12:43 sum(rate(lab_demo_requests_total[1m])) = 0 and the raw counter had plateaued. | Rate and error rate fell to 0. Counter flat. | up{job="otel-collector"} stayed 1. Prometheus was scraping the Collector, not the app. | Yes — you can tell “app gone” from “scrape target gone.” |
Point the SDK at the wrong port (127.0.0.1:4399) | Started 23:12:57. Eight GET /work calls returned HTTP 200. | SDK log: dial tcp 127.0.0.1:4399: connect: connection refused (first line 23:13:09). Collector debug exporter’s last metrics line stayed at 23:12:18 UTC (17:42:18Z). | Counter stuck at 473. Collector up still 1. The app looked healthy to curl. | Yes — green HTTP with a frozen counter means the pipe is broken, not the handler. |
| Stop the Collector, leave Prometheus | Killed 23:13:25. At 23:13:43 up{job="otel-collector"} = 0. | Scrape health drops. Grafana eventually has nothing new to draw for lab_demo_*. | up{job="prometheus"} stayed 1. | Yes — staleness of app series vs a down scrape target are different failures. |
histogram_quantile without by (le) | Queried while the series still existed. | histogram_quantile(0.99, sum(rate(lab_demo_latency_seconds_bucket[5m]))) returned an empty vector. | The correct query sum by (le) (rate(...[5m])) returned p99 = 571 ms over that wider window (it includes the slow ramp; the steady 2-minute window above was 49.9 ms). | Yes — a missing le grouping does not invent a pretty p99. It returns no series. How to read the 571 ms vs 49.9 ms gap is the <br/>percentiles post<br/>. |
If every panel stays green while you do this, you built decoration. Here the rate collapsed, the counter froze, and only the Collector-down step flipped up.
Honest advantages and disadvantages
| Approach | Advantages | Disadvantages |
|---|---|---|
| SDK → vendor directly | Fastest hello | Retries and filtering copied into every service |
| SDK → Collector → Prom → Grafana | Production-shaped; one place to batch and redact | More processes on a laptop (four, in this lab) |
| Metrics-only first | Useful dashboards the same day | Trace and log correlation waits |
| All three pillars on day one | Complete story | Easy to stop at green panels |
No “replace your vendor this weekend” claim. No hosted Grafana Cloud in this lab.
Footprint on this box (not a 30-minute soak)
RSS from ps during the short loaded run, a few minutes after start — not docker stats, and not a 30-minute idle measurement:
| Process | RSS |
|---|---|
| otelcol-contrib 0.111.0 | 155 MiB |
| Prometheus 2.54.1 (2 h retention, tiny scrape set) | 86 MiB (grew to ~86 MiB; ~78 MiB earlier in the run) |
| Grafana 11.2.0 | 170 MiB |
| Go demo | 16 MiB |
Together that is roughly 430 MiB before the OS and anything else. Do not co-host this stack with an app on a 512 MB VPS. Collector memory_limiter at 256 MiB is a backstop, not a sizing plan.
Same small-box theme, different runtime: Go GC + swap on a tiny VPS.
Related links:
FAQ
Do I need traces on day one?
No. Metrics that fail the right way are enough for this post. Add traces when you need a request story across services.
Collector contrib or core?
Core is smaller. Contrib has more exporters, including the Prometheus exporter used here. Pin the version either way — this lab pinned 0.111.0.
Why not scrape the app with Prometheus only?
You can, for a metrics-only process. The Collector is the point when you also want one pipe for traces and logs later. This lab teaches that pipe, metrics first.
Is remote write required?
No. Path A (scrape) was enough. Path C (OTLP into Prometheus) is a reasonable next experiment, not a requirement.
How does this relate to “average latency lies”?
This stack hosts the histogram. The percentiles playbook teaches which PromQL to trust. In this lab the steady 2-minute p99 was 49.9 ms and the 5-minute p99 was 571 ms — same metric, different window, different story.
Related links:
Can I run this on the same 512 MB VPS as my app?
Not this footprint. Collector + Prometheus + Grafana were ~430 MiB combined on a box that was not otherwise empty. Use a separate host, or drop Grafana and keep retention short, and measure before you recommend it.
More reading
Lab evidence
What I found running this
Host-binary lab 29 Sep 2026 IST (Docker unavailable). otelcol-contrib 0.111.0, Prometheus 2.54.1, Grafana 11.2.0, Go 1.24.4 demo. Steady: ~2.68 req/s, errors 0.105/s, p50/p95/p99 28.6/48.1/49.9 ms. Demo kill: rate 0 while collector up=1. Collector kill: otel up=0, Prometheus stayed up. histogram_quantile empty without by(le); with by(le) p99 571ms over 5m. RSS short run: collector 155MiB, prom 86, grafana 170, demo 16.
Related links
Plate 30
Why Your Average Latency Graph Is Lying (p50 / p95 / p99 Playbook)
Average latency hides tail pain. Learn when mean lies, how p50/p95/p99 work, why averaging quantiles fails, and how Prometheus histograms fix aggregation.
29 Sept 2026
Plate 42
Chrome DevTools AI Assistance (Gemini): Enable + Prompt Guide 2026
1 Oct 2026
Plate 46
Remix 3 Ships Oct 2, 2026: What Frontend Teams Need to Know
1 Oct 2026