ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 71

  1. Blog
  2. /Observability & SRE

zstd vs gzip vs lz4: Compression Throughput Lab

Hands-on zstd vs gzip vs lz4 on 64MiB payloads (real /usr text, mixed, urandom). Measured compress/decompress MB/s and ratio at default levels. No Docker.

Aditya Challa·30 September 2026·7 min read

Lab
On this page
  1. Intro — what this post promises
  2. What we are (and are not) measuring
  3. Lab topology
  4. Lead table — realistic text (text\_real, 64 MiB)
  5. Mixed (50% text + 50% urandom)
  6. Incompressible floor — os.urandom
  7. Pathological text\_rep (do not cite as typical)
  8. How to read these numbers
  9. Pitfalls we hit (or avoided)
  10. Practical checklist
  11. Methodology notes (reproducible)
  12. Verdict

Intro — what this post promises

“Just enable gzip” and “switch to zstd” are slide-deck sentences. Throughput and ratio depend on payload entropy and level, and pathological repeating text will flatter any codec.

This is a hands-on lab with measured numbers:

  1. Compress + decompress 64 MiB buffers in-process (Python bindings).
  2. Codecs: gzip (default 9 + level 6), zstd (default 3, level 1, level 19), lz4 frame default.
  3. Four payloads: pathological repeat, realistic /usr text, 50/50 mixed, os.urandom.
  4. Report ratio, compress MB/s, decompress MB/s (p50 wall).

Related links:

  • nginx gzip on vs off localhost lab
  • mmap vs read / O_DIRECT localhost lab
  • sendfile vs userspace copy localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • stdout buffering line vs full localhost lab

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5; python-zstandard 0.23.0 (CLI zstd 1.5.7); lz4 4.4.0+dfsg (CLI 1.10.0); system gzip 1.13. No Docker. Affiliates: 0. In-memory microbench — not HTTP Content-Encoding, not a threaded CLI bake-off.

Verdict up front: on realistic text, zstd default beat gzip9 by ~13× compress MB/s at a better ratio (2.39 vs 2.02). lz4 was the compress-speed king (~421 MB/s) with a weaker ratio (1.58). urandom stayed at ratio ≈1.000 for everyone.


What we are (and are not) measuring

LayerThis labSister labs
Codec CPU on a bufferyes—
nginx gzip on end-to-endno​nginx gzip lab​
Disk write of compressed bytesnofsync/fdatasync durable-write lab (packaged alongside; publish separately)
Wire capture / TLSnoopenssl TLS lab

Defaults used:

  • gzip.compress → level 9 (Python default); also level 6 as a common “fast gzip” knob.
  • zstandard.ZstdCompressor() → level 3; plus 1 and 19 for contrast.
  • lz4.frame.compress → library default frame format.

Related links:

  • nginx gzip on vs off localhost lab

Lab topology

64 MiB payloads in RAM → compress → decompress → assert round-trip
Payloads:
  text_rep   — short English line repeated (pathological)
  text_real  — concat unique text-ish files under /usr/share + Python stdlib
  mixed      — 50% text_real + 50% os.urandom
  random     — os.urandom 64 MiB
Repeats: 5 (zstd19: 2–3) after 1 warmup; MB/s from p50 wall

First draft used a “random” buffer built from repeating SHA-256 blocks — that still compressed ~80–100×. We reran with true os.urandom and a real multi-file corpus. Evidence notes call that out; quote the rerun.


Lead table — realistic text (text_real, 64 MiB)

CodecRatioCompress MB/sDecompress MB/sc p50
gzip9 (default)2.01517.81693.60 s
gzip62.01128.41642.25 s
zstd3 (default)2.3882276390.281 s
zstd11.9953656240.175 s
zstd192.7652.550725.3 s
lz4 frame1.5834216910.152 s

zstd3 vs gzip9: ~12.8× compress throughput, better ratio. lz4 wins raw compress speed here but gives up ~0.8 ratio points vs zstd3. zstd19 buys the best ratio (2.77) and loses two orders of magnitude on compress MB/s — offline archival vibes, not request path.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Mixed (50% text + 50% urandom)

CodecRatioc MB/sd MB/s
gzip91.30923.8253
gzip61.30832.8276
zstd31.3613221009
zstd11.301593774
zstd191.4073.1870
lz41.218518723

Half entropy halves the fairy tale. Ratios collapse toward ~1.2–1.4; zstd/lz4 still crush gzip on compress speed.


Incompressible floor — os.urandom

CodecRatioc MB/sd MB/s
gzip91.00040.7547
gzip61.00039.9559
zstd31.00011811721
zstd11.00014061698
zstd191.0002.91473
lz41.000535578

Nobody shrinks pure noise (compressed size ≈ raw; gzip slightly over 100% with framing). Speed here is “how fast can I give up” — zstd still scans fast; zstd19 does not.


Pathological text_rep (do not cite as typical)

Repeating one short line: gzip ~294×, lz4 ~243×, zstd ~10 800× at multi‑GB/s. Useful as an upper bound and a reminder that synthetic compressibility lies. Lead with text_real, not this arm.


How to read these numbers

  • Pick the payload class first — logs/HTML ≠ encrypted blobs ≠ /dev/urandom.
  • Default zstd was the balanced pick on realistic text (ratio + speed).
  • lz4 when CPU is dear and ratio can slip.
  • gzip9 remains the compatibility default; it was the slow compress path here.
  • Level 19 is a different product: ratio hobby, not RPS.

Related links:

  • nginx gzip on vs off localhost lab
  • process vs thread pool GIL localhost lab
  • asyncio vs threads IO concurrency lab

Pitfalls we hit (or avoided)

  1. Fake random from repeating digests — still highly compressible; fixed with os.urandom.
  2. Quoting pathological repeat ratios as “our logs.”
  3. Comparing CLI multi-thread zstd to single-buffer Python without labeling — this post is Python bindings, one buffer.
  4. Ignoring decompress — lz4/zstd both decompress fast here; gzip decompress was OK but compress was the bottleneck.
  5. Assuming level 19 is “free ratio” — 2.5 MB/s on text_real says otherwise.

Practical checklist

  • Measure on a slice of production-like bytes, not yes \| head.
  • Record ratio + compress MB/s + decompress MB/s at the level you will ship.
  • Prefer zstd default unless you need gzip clients or lz4 latency.
  • Keep level 19 / gzip9 for offline or compatibility lanes.
  • If the payload is already encrypted or random, skip compress — ratio ≈1.

Related links:

  • openssl TLS full vs resume localhost lab
  • nice / ionice CPU and disk priority lab
  • context-switch pipe ping-pong localhost lab


Methodology notes (reproducible)

All arms used in-process Python:

  • gzip.compress / gzip.decompress
  • zstandard.ZstdCompressor(level=…) / ZstdDecompressor
  • lz4.frame.compress / lz4.frame.decompress

Wall time is time.perf_counter around the full buffer op (not streaming chunk APIs). MB/s = 64 MiB / p50_seconds. Ratio = raw_bytes / compressed_bytes. Round-trip equality was asserted every repeat.

text_real was built by walking /usr/share/doc, /usr/share, and /usr/lib/python3.13, skipping ELF/gzip, deduping by a content hash prefix, concatenating until ≥64 MiB, then truncating. That is still “lab box text,” not your production JSON — but it is far closer than a single repeated sentence.

If you re-run on NVMe with the zstd CLI -T0, expect higher absolute MB/s; keep the relative ordering as the portable lesson unless you re-measure.

Related links:

  • epoll vs select FD_SETSIZE localhost lab
  • ulimit soft vs hard EMFILE lab

Verdict

On 64 MiB realistic /usr text, zstd level 3 delivered ratio 2.39 at ~227 MB/s compress and ~639 MB/s decompress — roughly 13× faster compress than gzip9 (2.02 / 17.8 / 169) with a better ratio. lz4 hit ~421 MB/s compress at ratio 1.58. urandom stayed at ≈1.000 ratio for all codecs. Choose the codec for the entropy you actually ship, and keep fairy-tale repeating buffers in the appendix.

Evidence path on the lab box: lab-evidence/33-zstd-vs-gzip-lz4/results/. Affiliates: 0.

zstd vs gziplz4 compressioncompression throughputzstandard levelgzip level 9localhost labsrepayload ratio

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; python-zstandard 0.23.0 (CLI zstd 1.5.7); lz4 4.4.0+dfsg (CLI 1.10.0); gzip 1.13. Payloads 64 MiB. Lead text_real: gzip9 ratio 2.02 c 17.8 MB/s d 169; gzip6 2.01 / 28.4 / 164; zstd3 2.39 / 227 / 639; zstd1 2.00 / 365 / 624; zstd19 2.77 / 2.5 / 507; lz4 1.58 / 421 / 691. urandom ratio ≈1.000 all; zstd3 ~1181 MB/s compress. Pathological text_rep labeled. No Docker. Affiliates: 0. Evidence: lab-evidence/33-zstd-vs-gzip-lz4/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we are (and are not) measuring
  3. Lab topology
  4. Lead table — realistic text (text\_real, 64 MiB)
  5. Mixed (50% text + 50% urandom)
  6. Incompressible floor — os.urandom
  7. Pathological text\_rep (do not cite as typical)
  8. How to read these numbers
  9. Pitfalls we hit (or avoided)
  10. Practical checklist
  11. Methodology notes (reproducible)
  12. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove