Plate 85
Nginx limit_req Rate-Limit Lab: Burst, nodelay, and the return Pitfall
Hands-on nginx 1.26 limit_req on localhost: 5r/s burst=10 nodelay yields 11/200 OK then 429s; return bypasses limiting; paced 5r/s stays green. Real numbers, no Docker.
Aditya Challa7 min read
On this page
- What limit\_req is (and is not)
- Lab topology
- Results (measured)
- Instant flood — 200 requests, 50 workers
- Paced clients on the limited port (`5r/s burst=10 nodelay`)
- Sanity: `1r/m burst=1 nodelay` sequential
- Pitfall: `return` bypasses `limit_req`
- What we would put in front of a real API next
- FAQ
- Why 429 instead of 503?
- Does `burst=10` mean 10 r/s?
- What about `limit_conn`?
- Will this stop distributed scrapers?
- Can I lab this without Docker?
- Verdict
- Related ShopperCove posts
Intro — what this post promises
Rate limiting is the cheapest “please stop hitting me” lever on a reverse proxy. Nginx’s limit_req is also one of the easiest knobs to misconfigure: the docs look short, the zone looks fine, and every request still returns 200.
This post is a localhost lab with measured numbers:
- What
limit_req_zone+limit_reqactually do (leaky bucket, not a hard RPS ceiling you can eyeball). - A baseline flood with no limit.
- The same flood with
rate=5r/s burst=10 nodelayand with no burst. - Paced clients at under / at / over the configured rate.
- The pitfall that made our first configs look “broken”:
returnin the samelocationbypasseslimit_req. - How this ties to latency graphs and a first observability stack — without repeating those posts.
Related links:
- Why your average latency graph is lying (p50 / p95 / p99)
- OpenTelemetry Collector + Prometheus + Grafana first stack
- How to read server monitoring graphs
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores). nginx/1.26.3 as a host binary (Debian package). Four separate nginx masters, each bound to 127.0.0.1 only on ports 18080–18083. No Docker. No public bind. Static try_files → index.html (“ok”). Client floods: Python urllib thread pool + paced sequential loops. Affiliates: 0.
Verdict up front: limit_req works and is worth running on login, search, and other abuse-prone paths — but only after you confirm 429s under a real burst, not after a green config test.
What limit_req is (and is not)
Nginx documents limit_req_zone as a shared-memory zone keyed by a variable (usually $binary_remote_addr) with a rate, and limit_req as the per-location application of that zone with an optional burst and nodelay.
Related links:
Mental model that matched our lab:
| Knob | What it did here |
|---|---|
rate=5r/s | Sustained refill ~5 requests/second for that key |
burst=10 | Allows a short queue/excess of 10 above the rate |
nodelay | Excess burst requests are not delayed — they succeed immediately until burst is spent, then reject |
no burst | Instantaneous flood collapses to one success, then 429s |
limit_req_status 429 | Reject with 429 instead of nginx’s default 503 |
This is not a global “API quota” product. It is per-key (here: one IP). Behind a CDN you need real-IP / trusted proxy setup or you rate-limit the CDN edge, not the client.
Lab topology
Minimal limited location (shape we ended on):
We deliberately did not put return 200 "ok\n"; in that location. See the pitfall section.
Results (measured)
Instant flood — 200 requests, 50 workers
| Config | Wall | Client RPS | HTTP 200 | HTTP 429 | OK % |
|---|---|---|---|---|---|
| Baseline (no limit) | 0.091 s | ~2210 | 200 | 0 | 100% |
5r/s + burst=10 nodelay | 0.105 s | ~1903 | 11 | 189 | 5.5% |
5r/s + no burst | 0.105 s | ~1914 | 1 | 199 | 0.5% |
The 11 successes with burst=10 match the expected “about one in-rate token + 10 burst” story for an instantaneous dump against an empty bucket. The 1 success with no burst matches “empty bucket allows the first request, rejects the rest.”
Repeat flood on the limited port (100 requests / 40 workers after refill sleep): again 11×200 and 89×429. Same shape, not a fluke.
Error log (warn level) on rejects looked like:
Paced clients on the limited port (5r/s burst=10 nodelay)
| Pace | N | Wall | 200 | 429 | OK % |
|---|---|---|---|---|---|
| 5 r/s (at config) | 30 | 6.003 s | 30 | 0 | 100% |
| 20 r/s (4× over) | 40 | 2.004 s | 20 | 20 | 50% |
| 2 r/s (under) | 20 | 10.002 s | 20 | 0 | 100% |
At the configured rate, the leaky bucket keeps up. Well under the rate, everything is green. At 20 r/s, burst absorbs some of the early excess, then half the sample is rejected — which is exactly the “burst is not infinite headroom” lesson.
Sanity: 1r/m burst=1 nodelay sequential
10 quick sequential GETs: 2×200, then 8×429. Confirms rejecting works even without a thread pool.
Pitfall: return bypasses limit_req
Our first configs used:
Config test passed. Floods returned 200/200. Error log stayed quiet. It looked like limit_req was broken.
Nginx phase order is the reason: return is handled in the rewrite phase; limit_req runs in preaccess. Rewrite finishes first, so the request never reaches the limiter.
Fix that worked in this lab: serve a static file with root + try_files (content phase), or proxy_pass an upstream. Then the same limit_req line produced the 429 tables above.
If your “rate limit” never fires in staging, check for return / early rewrite exits in the same location before you tune burst.
What we would put in front of a real API next
- Bind carefully — this lab used
127.0.0.1only. Production listens where the load balancer expects, with TLS at the edge. - Pick the key — IP is the default; authenticated routes often want a user id (and a separate, stricter zone for anonymous).
- Status code — set
limit_req_status 429so clients and scrapers see Too Many Requests, not a fake upstream failure (503). - Observe rejects — count 429s and
limiting requestswarnings next to golden signals. Rate-limit mistunes show up as traffic and error-rate symptoms, not as “CPU is fine.” - Do not confuse with latency SLOs — a healthy limited path still needs p95/p99 on the 200 path; rejecting junk should not be your only latency strategy.
Related links:
- Why your average latency graph is lying
- OTel + Prometheus + Grafana first stack
- How to read server monitoring graphs
FAQ
Why 429 instead of 503?
Nginx defaults limit_req rejects to 503. We set limit_req_status 429 so the client semantics match “slow down,” not “server broken.” Either works operationally; be consistent with what your clients and alerts expect.
Does burst=10 mean 10 r/s?
No. rate is the sustained refill. burst is short-term excess capacity. In our instant flood, burst=10 meant about 11 successes, not 10 r/s forever.
What about limit_conn?
Different tool: connection concurrency, not request rate. Useful for slowloris-style pressure; not a substitute for limit_req on chatty HTTP/1.1 or HTTP/2 request storms.
Will this stop distributed scrapers?
Per-IP limits stop noisy single sources. Distributed scrapers need edge/WAF/bot management and application auth. limit_req is still useful as a cheap inner fuse.
Can I lab this without Docker?
Yes. This entire post ran on a host nginx binary and localhost ports.
Verdict
Ship limit_req on abuse-prone locations after a burst test that proves 429s. Our numbers: baseline flood 200/200 OK; with 5r/s burst=10 nodelay, the same flood dropped to 11 OK and 189×429; with no burst, 1 OK and 199×429; paced traffic at 5 r/s stayed 100% OK. The config that “passes nginx -t” is not enough — especially if return short-circuits the limiter.
Affiliates: 0. No product hop. Evidence lives under the lab tree used for this write-up (lab-evidence/11-nginx-rate-limit/).
Related ShopperCove posts
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. nginx/1.26.3 host binary, localhost-only :18080–18083, try_files static OK. Baseline 200/50 workers: 200×200 in 0.091s (~2209 client rps). limit_req rate=5r/s burst=10 nodelay: 11×200 + 189×429 (5.5% OK). No burst: 1×200 + 199×429. Paced 5r/s×30: 100% 200; paced 20r/s×40: 20×200 + 20×429; paced 2r/s×20: 100% 200. Repeat burst 100/40: again 11×200. Pitfall: location return runs rewrite before limit_req preaccess — all 200s until switched to try_files. limit_req_status 429. Affiliates: 0. Evidence: lab-evidence/11-nginx-rate-limit/.
Related links
Plate 42
Chrome DevTools AI Assistance (Gemini): Enable + Prompt Guide 2026
1 Oct 2026
Plate 69
Next.js 16 App Router Production Checklist (2026)
A practical Next.js 16 App Router production checklist for Server Components, PPR, caching, streaming, metadata, and SEO.
1 Oct 2026
Plate 22
Nginx Gzip On vs Off: Localhost Wire Size and RPS Lab
Hands-on nginx gzip on/off lab: HTML ~107x smaller at level 1; JSON RPS ~1987 off vs ~1031 l1 vs ~563 l6 on localhost. Measured numbers, no Docker.
30 Sept 2026