ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 30

  1. Blog
  2. /Observability & SRE

Stdout Buffering Lab: Line vs Full vs Unbuffered (Real Numbers)

Hands-on stdout buffering lab: full buffering maximizes throughput, while line and unbuffered modes restore live visibility on pipes.

Aditya Challa·30 September 2026·5 min read

Lab
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — pipe throughput (200k lines, p50)
  5. Arm B — visibility latency (40 lines × 5 ms, p50)
  6. Arm C — file redirect (200k lines, p50)
  7. Arm D — default `print()` vs `-u` / env
  8. When full buffering still wins
  9. Pitfalls we hit (or avoided)
  10. Practical checklist
  11. Verdict

Intro — what this post promises

Your job “hangs” with no logs — then dumps everything at exit. That is often stdout buffering, not a stuck process. Line buffering vs full block buffering is the difference between seeing progress and staring at a quiet pipe.

This is a hands-on lab with measured numbers:

  1. Throughput for 200k lines under full (8 KiB), line (buffering=1), and unbuffered (buffering=0) modes.
  2. First-line visibility when a child prints 40 lines at 5 ms pace into a pipe.
  3. File-redirect throughput (same modes).
  4. Default print() on a pipe vs python -u vs PYTHONUNBUFFERED=1.

Related links:

  • Pipe vs tmpfile IPC localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • Nginx gzip on vs off localhost lab
  • How to read server monitoring graphs

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. Child stdout is a pipe or file (not a TTY — so Python’s default is block-buffered). No Docker. No GPU. No API keys. Affiliates: 0.

Verdict up front: full buffering hit ~4.0M lines/s on a pipe but first line stayed invisible until ~226 ms. Line / unbuffered showed first line at ~12 ms. Default pipe print() matched full; python -u and PYTHONUNBUFFERED=1 restored early visibility.


What we compared

ModeHow we opened stdoutTypical use
fulltext, buffering=8192high-volume logs to pipe/file
linetext, buffering=1progress lines you want live
nonebinary, buffering=0every write hits the OS

Related links:

  • TextIOWrapper buffering — Python docs
  • python -u / PYTHONUNBUFFERED

Lab topology

Parent reads child stdout (pipe) or redirects to a tempfile
Child: write N lines under full | line | none, then DONE + flush
Arm A: 200k lines throughput (pipe)
Arm B: 40 lines @ 5 ms — first / mid / done visibility
Arm C: 200k lines → file
Arm D: default print() vs python -u vs PYTHONUNBUFFERED=1

Arm A — pipe throughput (200k lines, p50)

ModeWall sLines/s
full0.050~4.03M
line0.2430.82M (4.9× slower)
none0.158~1.27M

Block buffering wins raw throughput. Paying a write(2) per line (or nearly) is expensive when you do not need live visibility.


Arm B — visibility latency (40 lines × 5 ms, p50)

ModeFirst line sMid sDONE s
full0.2260.2270.227
line0.0120.1190.226
none0.0120.1190.229

Full buffering held everything until the final flush — first, mid, and DONE arrived together near ~226 ms. Line and unbuffered streamed: first line ~12 ms after start.


Arm C — file redirect (200k lines, p50)

ModeWall sLines/s
full0.047~4.30M
line0.3540.57M (7.6× slower)
none0.352~0.57M

Same story on a file sink: full buffering dominates; line and none collapse together when every line is a syscall.


Arm D — default print() vs -u / env

SetupFirst line sDONE s
default pipe0.2290.229
python -u0.0140.228
PYTHONUNBUFFERED=10.0130.228

Default print() to a pipe behaved like full buffering. Both -u and PYTHONUNBUFFERED=1 restored early first-line visibility (~13–14 ms).

Related links:

  • Pipe vs tmpfile IPC localhost lab

When full buffering still wins

  • High-volume batch logs where you only care about the finished file/pipe contents.
  • CI artifacts / redirects where progress is not consumed live.
  • Throughput-sensitive emitters (Arm A/C: ~4–5×–7× faster than line mode here).

Use line buffering (or -u / explicit flush) when a human, supervisor, or sidecar reads the pipe live.


Pitfalls we hit (or avoided)

  1. Assuming TTY rules apply to pipes — non-TTY stdout is block-buffered by default.
  2. Blaming “hung” jobs when logs are just sitting in the 8 KiB buffer.
  3. Line-buffering every hot path — Arm A/C show the throughput tax.
  4. Forgetting the final flush — full mode only became visible after flush() / exit.
  5. Mixing parent bufsize with child buffering — we set parent pipe reads unbuffered so visibility reflects the child.

Practical checklist

  • Live progress on pipes: python -u, PYTHONUNBUFFERED=1, or explicit flush=True / line buffering.
  • Bulk redirects: prefer default/full buffering; do not pay per-line syscalls.
  • Supervisors/sidecar log shippers: test with a pipe, not only an interactive TTY.
  • Report sink type (TTY/pipe/file) + mode with every lines/s claim.
  • Pair with pipe-vs-tmpfile when the question is IPC shape, not print buffering.

Verdict

On non-TTY stdout, full (8 KiB) buffering delivered ~4.0M lines/s (pipe) / ~4.3M (file) but hid the first line until ~226 ms. Line / unbuffered showed first line at ~12 ms and paid ~5–8× throughput. Default pipe print() matched full; python -u and PYTHONUNBUFFERED=1 fixed visibility without rewriting the child.

Evidence path on the lab box: lab-evidence/25-stdout-buffering/results/. Affiliates: 0.

stdout bufferingpython -upythonunbufferedline buffered vs block bufferedpipe log visibilitylocalhost labsreprint flush

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5. Non-TTY child stdout. Pipe 200k lines p50: full 0.050s (~4.03M lines/s); line 0.243s (~0.82M); none 0.158s (1.27M). Visibility 40×5ms: full first=0.226s (batch); line/none first0.012s. File 200k: full 0.047s (~4.30M); line/none ~0.35s (~0.57M). Default pipe print first=0.229s; python -u 0.014s; PYTHONUNBUFFERED=1 0.013s. No Docker. Affiliates: 0. Evidence: lab-evidence/25-stdout-buffering/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — pipe throughput (200k lines, p50)
  5. Arm B — visibility latency (40 lines × 5 ms, p50)
  6. Arm C — file redirect (200k lines, p50)
  7. Arm D — default `print()` vs `-u` / env
  8. When full buffering still wins
  9. Pitfalls we hit (or avoided)
  10. Practical checklist
  11. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove