ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 44

  1. Blog

OpenTelemetry Collector + Prometheus + Grafana: A First Stack That Actually Diagnoses

A first OTel Collector, Prometheus, and Grafana lab: OTLP in, Prometheus scrape out, and a break-one-signal drill with measured PromQL.

Aditya Challa·29 September 2026·10 min read

Summary
On this page
  1. Intro — what this post promises
  2. What each piece owns
  3. Topology
  4. Collector config — the pieces that mattered
  5. Three ways metrics reach Prometheus — we ran Path A
  6. Grafana — the minimum dashboard that diagnosed
  7. What the demo emitted
  8. The break-one-signal drill
  9. Honest advantages and disadvantages
  10. Footprint on this box (not a 30-minute soak)
  11. FAQ
  12. More reading

Intro — what this post promises

Most “OTel + Prometheus + Grafana” tutorials stop when the dashboard turns green. The bar here is a stack that diagnoses: when the app dies, when the SDK dials the wrong Collector port, when the Collector itself stops — which panel moves, and which one stays calm on purpose?

This is the opinionated minimal lab for that question:

  1. Why a Collector belongs in the path.
  2. The topology (app → OTLP → Collector → Prometheus → Grafana).
  3. One metrics path, run for real: Collector Prometheus exporter, scraped.
  4. A break-one-signal drill with timestamps from 29 Sep 2026 (IST).
  5. Where percentiles and golden signals live — other posts, not this one.

Golden-signal literacy stays on How to read server monitoring graphs. Histogram and quantile rules stay on Why your average latency graph is lying (CMS auto-slug until the slug editor ships). This post only hosts the panels those playbooks query.

Related links:

  • How to read server monitoring graphs
  • Why your average latency graph is lying

Lab honesty: Docker was not installed on the lab box (docker not found). The same pinned versions ran as host binaries, not containers: OpenTelemetry Collector contrib 0.111.0, Prometheus 2.54.1, Grafana 11.2.0, and a tiny Go 1.24.4 demo using the OpenTelemetry Go SDK v1.32.0 (otlpmetricgrpc v1.32.0). Path used: A — prometheus exporter + scrape. otelcol-contrib validate exited 0. This was a few minutes of load on a shared Linux box, not a 30-minute idle soak and not a production cluster.

A practical first OpenTelemetry stack is: app → OTLP → Collector → Prometheus → Grafana. The Collector receives OTLP, batches, and exposes a Prometheus text endpoint. Grafana queries Prometheus. Prove the stack by breaking one signal and watching the right pillar fail.


What each piece owns

ComponentJob in this labWhere to start
App / SDKEmit metrics via OTLP/gRPCLanguage OTel SDK docs
OpenTelemetry CollectorReceive, batch, exportCollector docs
PrometheusStore and query metrics (PromQL)Prometheus scrape config
GrafanaDashboards on top of PrometheusPrometheus data source

OpenTelemetry’s Collector overview recommends a Collector next to services in real deployments: apps offload quickly; the Collector handles retries, batching, and filtering. Direct SDK → backend is fine for a tiny experiment. This lab still uses a Collector so the shape matches production.

Related links:

  • OpenTelemetry’s Collector overview

Topology

[Go demo :8080] --OTLP/gRPC 127.0.0.1:4317--> [otelcol-contrib 0.111.0]
                                                    | metrics pipeline
                                                    v
                                          prometheus exporter :8889
                                                    ^
                                          scrape every 5s (Path A)
                                                    |
                                          [Prometheus 2.54.1 :9090]
                                                    ^
                                                    | PromQL
                                          [Grafana 11.2.0 :3000]

Traces and logs are out of scope. Ship metrics first so the percentiles playbook has a real histogram to query.

Related links:

  • percentiles playbook

Listeners in this lab were bound to 127.0.0.1, not 0.0.0.0. Collector docs warn that an open OTLP port is a denial-of-service surface. On a laptop that is the right default; on a VPS, firewall the port or keep it on a private interface.


Collector config — the pieces that mattered

Collector config is always the same spine (configuration guide): receivers → processors → exporters, then service.pipelines that actually enable them. A component that is not listed in a pipeline does nothing.

Related links:

  • configuration guide

This is the file that validated and ran (contrib 0.111.0):

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 127.0.0.1:4317
      http:
        endpoint: 127.0.0.1:4318

processors:
  memory_limiter:
    check_interval: 5s
    limit_mib: 256
    spike_limit_mib: 64
  batch:
    timeout: 2s
    send_batch_size: 512

exporters:
  prometheus:
    endpoint: 127.0.0.1:8889
    namespace: lab
    resource_to_telemetry_conversion:
      enabled: true
  debug:
    verbosity: basic

service:
  pipelines:
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [prometheus, debug]
otelcol-contrib validate --config=otel-config.yaml
# exit 0 on contrib 0.111.0

namespace: lab is why the demo’s demo.requests counter showed up in Prometheus as lab_demo_requests_total. Resource attributes (service.name, deployment.environment) were copied onto the series because resource_to_telemetry_conversion was on.


Three ways metrics reach Prometheus — we ran Path A

Interoperability is documented from both sides (OTel “Prom and OTel”, OTLP metrics export to Prometheus):

Related links:

  • OTel “Prom and OTel”
  • OTLP metrics export to Prometheus
PathHow it worksThis lab?
A. prometheus exporter + scrapeCollector serves Prom text; Prometheus scrapes :8889Yes — this is what ran
B. prometheusremotewriteCollector pushes to Prom’s remote-write receiverNot run
C. OTLP → PrometheusPush OTLP HTTP to Prom’s native OTLP metrics endpointNot run

Path A is the right first lab because the mental model is still “Prometheus scrapes a target.” Path C is the more OTel-native production story on modern Prometheus; do not implement all three in one afternoon.

Prometheus config that scraped the Collector (interval 5s):

global:
  scrape_interval: 5s

scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets: ["127.0.0.1:9090"]
  - job_name: otel-collector
    metrics_path: /metrics
    static_configs:
      - targets: ["127.0.0.1:8889"]

Working PromQL at 23:10 IST (steady demo load, 2-minute rate window):

QueryResult
up{job="otel-collector"}1
sum(rate(lab_demo_requests_total[1m]))2.68 req/s
sum(rate(lab_demo_errors_total[1m]))0.105 req/s
histogram_quantile(0.50, sum by (le) (rate(lab_demo_latency_seconds_bucket[2m])))28.6 ms
same, φ=0.9548.1 ms
same, φ=0.9949.9 ms
lab_demo_requests_total169 (HTTP 200) and 7 (HTTP 500) at that instant

The cumulative counter later froze at 473 when the demo stopped (see the drill). Series carried service_name="shoppercove-demo", deployment_environment="lab", and http_route="/work".

If you would rather ship containers, pin the same versions. This Compose file was not executed here (no Docker). It matches the binaries that were:

services:
  otel-collector:
    image: otel/opentelemetry-collector-contrib:0.111.0
    command: ["--config=/etc/otelcol/config.yaml"]
    volumes: ["./otel-config.yaml:/etc/otelcol/config.yaml:ro"]
    ports: ["4317:4317", "4318:4318", "8889:8889"]
  prometheus:
    image: prom/prometheus:v2.54.1
    volumes: ["./prometheus.yml:/etc/prometheus/prometheus.yml:ro"]
    ports: ["9090:9090"]
  grafana:
    image: grafana/grafana:11.2.0
    ports: ["3000:3000"]

Grafana — the minimum dashboard that diagnosed

Prometheus data source: http://127.0.0.1:9090 (inside Compose that URL is http://prometheus:9090). Grafana 11.2.0 provisioned five panels:

  1. Request rate — sum(rate(lab_demo_requests_total[1m]))
  2. Error rate — sum(rate(lab_demo_errors_total[1m]))
  3. Latency quantiles — histogram_quantile on lab_demo_latency_seconds_bucket for p50 / p95 / p99. Rules for when those numbers lie are in the percentiles post, not here.
  4. Scrape health — up{job=~"otel-collector|prometheus"}
  5. Raw counter — lab_demo_requests_total (the flatline is easier to trust than a rate window)

Related links:

  • percentiles post

From process start (23:08 IST) to a Grafana panel that showed request rate (23:11 IST) was about two minutes. The Collector’s :8889/metrics already listed lab_demo_* about 45 seconds after the demo started.


What the demo emitted

One Go process, not a microservice estate. OpenTelemetry Go SDK v1.32.0, exporter otlpmetricgrpc v1.32.0, gRPC v1.67.1, runtime Go 1.24.4.

SignalOTel namePrometheus name (namespace lab)
Request counterdemo.requestslab_demo_requests_total
Error counterdemo.errorslab_demo_errors_total
Latency histogramdemo.latency (unit seconds, explicit buckets 5 ms … 10 s)lab_demo_latency_seconds_bucket
Resourceservice.name=shoppercove-demo, deployment.environment=lablabels on the series

Export interval was 2 seconds to 127.0.0.1:4317. Proof the series existed: the PromQL table above, and GET http://127.0.0.1:8889/metrics returning lab_demo_requests_total{http_status_code="200",...}.


The break-one-signal drill

All times 29 Sep 2026, Asia/Calcutta. Scrape interval 5 seconds.

BreakWhen (IST)What movedWhat stayedPass?
Stop the demo processKilled 23:12:18. At 23:12:43 sum(rate(lab_demo_requests_total[1m])) = 0 and the raw counter had plateaued.Rate and error rate fell to 0. Counter flat.up{job="otel-collector"} stayed 1. Prometheus was scraping the Collector, not the app.Yes — you can tell “app gone” from “scrape target gone.”
Point the SDK at the wrong port (127.0.0.1:4399)Started 23:12:57. Eight GET /work calls returned HTTP 200.SDK log: dial tcp 127.0.0.1:4399: connect: connection refused (first line 23:13:09). Collector debug exporter’s last metrics line stayed at 23:12:18 UTC (17:42:18Z).Counter stuck at 473. Collector up still 1. The app looked healthy to curl.Yes — green HTTP with a frozen counter means the pipe is broken, not the handler.
Stop the Collector, leave PrometheusKilled 23:13:25. At 23:13:43 up{job="otel-collector"} = 0.Scrape health drops. Grafana eventually has nothing new to draw for lab_demo_*.up{job="prometheus"} stayed 1.Yes — staleness of app series vs a down scrape target are different failures.
histogram_quantile without by (le)Queried while the series still existed.histogram_quantile(0.99, sum(rate(lab_demo_latency_seconds_bucket[5m]))) returned an empty vector.The correct query sum by (le) (rate(...[5m])) returned p99 = 571 ms over that wider window (it includes the slow ramp; the steady 2-minute window above was 49.9 ms).Yes — a missing le grouping does not invent a pretty p99. It returns no series. How to read the 571 ms vs 49.9 ms gap is the <br/>percentiles post<br/>.

If every panel stays green while you do this, you built decoration. Here the rate collapsed, the counter froze, and only the Collector-down step flipped up.


Honest advantages and disadvantages

ApproachAdvantagesDisadvantages
SDK → vendor directlyFastest helloRetries and filtering copied into every service
SDK → Collector → Prom → GrafanaProduction-shaped; one place to batch and redactMore processes on a laptop (four, in this lab)
Metrics-only firstUseful dashboards the same dayTrace and log correlation waits
All three pillars on day oneComplete storyEasy to stop at green panels

No “replace your vendor this weekend” claim. No hosted Grafana Cloud in this lab.


Footprint on this box (not a 30-minute soak)

RSS from ps during the short loaded run, a few minutes after start — not docker stats, and not a 30-minute idle measurement:

ProcessRSS
otelcol-contrib 0.111.0155 MiB
Prometheus 2.54.1 (2 h retention, tiny scrape set)86 MiB (grew to ~86 MiB; ~78 MiB earlier in the run)
Grafana 11.2.0170 MiB
Go demo16 MiB

Together that is roughly 430 MiB before the OS and anything else. Do not co-host this stack with an app on a 512 MB VPS. Collector memory_limiter at 256 MiB is a backstop, not a sizing plan.

Same small-box theme, different runtime: Go GC + swap on a tiny VPS.

Related links:

  • Go GC + swap on a tiny VPS

FAQ

Do I need traces on day one?
No. Metrics that fail the right way are enough for this post. Add traces when you need a request story across services.

Collector contrib or core?
Core is smaller. Contrib has more exporters, including the Prometheus exporter used here. Pin the version either way — this lab pinned 0.111.0.

Why not scrape the app with Prometheus only?
You can, for a metrics-only process. The Collector is the point when you also want one pipe for traces and logs later. This lab teaches that pipe, metrics first.

Is remote write required?
No. Path A (scrape) was enough. Path C (OTLP into Prometheus) is a reasonable next experiment, not a requirement.

How does this relate to “average latency lies”?
This stack hosts the histogram. The percentiles playbook teaches which PromQL to trust. In this lab the steady 2-minute p99 was 49.9 ms and the 5-minute p99 was 571 ms — same metric, different window, different story.

Related links:

  • percentiles playbook

Can I run this on the same 512 MB VPS as my app?
Not this footprint. Collector + Prometheus + Grafana were ~430 MiB combined on a box that was not otherwise empty. Use a separate host, or drop Grafana and keep retention short, and measure before you recommend it.


More reading

  • How to read server monitoring graphs
  • Why your average latency graph is lying
  • Go GC + swap on a tiny VPS
  • Collector configuration
  • Prometheus histograms and summaries
  • OTel “Prom and OTel”
  • OTLP metrics export to Prometheus
opentelemetry collectorprometheusgrafanaotlppromqlhistogram_quantilefirst observability stack

Lab evidence

What I found running this

Host-binary lab 29 Sep 2026 IST (Docker unavailable). otelcol-contrib 0.111.0, Prometheus 2.54.1, Grafana 11.2.0, Go 1.24.4 demo. Steady: ~2.68 req/s, errors 0.105/s, p50/p95/p99 28.6/48.1/49.9 ms. Demo kill: rate 0 while collector up=1. Collector kill: otel up=0, Prometheus stayed up. histogram_quantile empty without by(le); with by(le) p99 571ms over 5m. RSS short run: collector 155MiB, prom 86, grafana 170, demo 16.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 30

    Why Your Average Latency Graph Is Lying (p50 / p95 / p99 Playbook)

    Average latency hides tail pain. Learn when mean lies, how p50/p95/p99 work, why averaging quantiles fails, and how Prometheus histograms fix aggregation.

    29 Sept 2026

  • Plate 42

    Chrome DevTools AI Assistance (Gemini): Enable + Prompt Guide 2026

    1 Oct 2026

  • Plate 46

    Remix 3 Ships Oct 2, 2026: What Frontend Teams Need to Know

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What each piece owns
  3. Topology
  4. Collector config — the pieces that mattered
  5. Three ways metrics reach Prometheus — we ran Path A
  6. Grafana — the minimum dashboard that diagnosed
  7. What the demo emitted
  8. The break-one-signal drill
  9. Honest advantages and disadvantages
  10. Footprint on this box (not a 30-minute soak)
  11. FAQ
  12. More reading
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove