Plate 71
zstd vs gzip vs lz4: Compression Throughput Lab
Hands-on zstd vs gzip vs lz4 on 64MiB payloads (real /usr text, mixed, urandom). Measured compress/decompress MB/s and ratio at default levels. No Docker.
Aditya Challa7 min read
On this page
- Intro — what this post promises
- What we are (and are not) measuring
- Lab topology
- Lead table — realistic text (text\_real, 64 MiB)
- Mixed (50% text + 50% urandom)
- Incompressible floor — os.urandom
- Pathological text\_rep (do not cite as typical)
- How to read these numbers
- Pitfalls we hit (or avoided)
- Practical checklist
- Methodology notes (reproducible)
- Verdict
Intro — what this post promises
“Just enable gzip” and “switch to zstd” are slide-deck sentences. Throughput and ratio depend on payload entropy and level, and pathological repeating text will flatter any codec.
This is a hands-on lab with measured numbers:
- Compress + decompress 64 MiB buffers in-process (Python bindings).
- Codecs: gzip (default 9 + level 6), zstd (default 3, level 1, level 19), lz4 frame default.
- Four payloads: pathological repeat, realistic /usr text, 50/50 mixed,
os.urandom. - Report ratio, compress MB/s, decompress MB/s (p50 wall).
Related links:
- nginx gzip on vs off localhost lab
- mmap vs read / O_DIRECT localhost lab
- sendfile vs userspace copy localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
- stdout buffering line vs full localhost lab
Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5; python-zstandard 0.23.0 (CLI zstd 1.5.7); lz4 4.4.0+dfsg (CLI 1.10.0); system gzip 1.13. No Docker. Affiliates: 0. In-memory microbench — not HTTP Content-Encoding, not a threaded CLI bake-off.
Verdict up front: on realistic text, zstd default beat gzip9 by ~13× compress MB/s at a better ratio (2.39 vs 2.02). lz4 was the compress-speed king (~421 MB/s) with a weaker ratio (1.58). urandom stayed at ratio ≈1.000 for everyone.
What we are (and are not) measuring
| Layer | This lab | Sister labs |
|---|---|---|
| Codec CPU on a buffer | yes | — |
nginx gzip on end-to-end | no | nginx gzip lab |
| Disk write of compressed bytes | no | fsync/fdatasync durable-write lab (packaged alongside; publish separately) |
| Wire capture / TLS | no | openssl TLS lab |
Defaults used:
gzip.compress→ level 9 (Python default); also level 6 as a common “fast gzip” knob.zstandard.ZstdCompressor()→ level 3; plus 1 and 19 for contrast.lz4.frame.compress→ library default frame format.
Related links:
Lab topology
First draft used a “random” buffer built from repeating SHA-256 blocks — that still compressed ~80–100×. We reran with true os.urandom and a real multi-file corpus. Evidence notes call that out; quote the rerun.
Lead table — realistic text (text_real, 64 MiB)
| Codec | Ratio | Compress MB/s | Decompress MB/s | c p50 |
|---|---|---|---|---|
| gzip9 (default) | 2.015 | 17.8 | 169 | 3.60 s |
| gzip6 | 2.011 | 28.4 | 164 | 2.25 s |
| zstd3 (default) | 2.388 | 227 | 639 | 0.281 s |
| zstd1 | 1.995 | 365 | 624 | 0.175 s |
| zstd19 | 2.765 | 2.5 | 507 | 25.3 s |
| lz4 frame | 1.583 | 421 | 691 | 0.152 s |
zstd3 vs gzip9: ~12.8× compress throughput, better ratio. lz4 wins raw compress speed here but gives up ~0.8 ratio points vs zstd3. zstd19 buys the best ratio (2.77) and loses two orders of magnitude on compress MB/s — offline archival vibes, not request path.
Related links:
Mixed (50% text + 50% urandom)
| Codec | Ratio | c MB/s | d MB/s |
|---|---|---|---|
| gzip9 | 1.309 | 23.8 | 253 |
| gzip6 | 1.308 | 32.8 | 276 |
| zstd3 | 1.361 | 322 | 1009 |
| zstd1 | 1.301 | 593 | 774 |
| zstd19 | 1.407 | 3.1 | 870 |
| lz4 | 1.218 | 518 | 723 |
Half entropy halves the fairy tale. Ratios collapse toward ~1.2–1.4; zstd/lz4 still crush gzip on compress speed.
Incompressible floor — os.urandom
| Codec | Ratio | c MB/s | d MB/s |
|---|---|---|---|
| gzip9 | 1.000 | 40.7 | 547 |
| gzip6 | 1.000 | 39.9 | 559 |
| zstd3 | 1.000 | 1181 | 1721 |
| zstd1 | 1.000 | 1406 | 1698 |
| zstd19 | 1.000 | 2.9 | 1473 |
| lz4 | 1.000 | 535 | 578 |
Nobody shrinks pure noise (compressed size ≈ raw; gzip slightly over 100% with framing). Speed here is “how fast can I give up” — zstd still scans fast; zstd19 does not.
Pathological text_rep (do not cite as typical)
Repeating one short line: gzip ~294×, lz4 ~243×, zstd ~10 800× at multi‑GB/s. Useful as an upper bound and a reminder that synthetic compressibility lies. Lead with text_real, not this arm.
How to read these numbers
- Pick the payload class first — logs/HTML ≠ encrypted blobs ≠
/dev/urandom. - Default zstd was the balanced pick on realistic text (ratio + speed).
- lz4 when CPU is dear and ratio can slip.
- gzip9 remains the compatibility default; it was the slow compress path here.
- Level 19 is a different product: ratio hobby, not RPS.
Related links:
- nginx gzip on vs off localhost lab
- process vs thread pool GIL localhost lab
- asyncio vs threads IO concurrency lab
Pitfalls we hit (or avoided)
- Fake random from repeating digests — still highly compressible; fixed with
os.urandom. - Quoting pathological repeat ratios as “our logs.”
- Comparing CLI multi-thread zstd to single-buffer Python without labeling — this post is Python bindings, one buffer.
- Ignoring decompress — lz4/zstd both decompress fast here; gzip decompress was OK but compress was the bottleneck.
- Assuming level 19 is “free ratio” — 2.5 MB/s on text_real says otherwise.
Practical checklist
- Measure on a slice of production-like bytes, not
yes \| head. - Record ratio + compress MB/s + decompress MB/s at the level you will ship.
- Prefer zstd default unless you need gzip clients or lz4 latency.
- Keep level 19 / gzip9 for offline or compatibility lanes.
- If the payload is already encrypted or random, skip compress — ratio ≈1.
Related links:
- openssl TLS full vs resume localhost lab
- nice / ionice CPU and disk priority lab
- context-switch pipe ping-pong localhost lab
Methodology notes (reproducible)
All arms used in-process Python:
gzip.compress/gzip.decompresszstandard.ZstdCompressor(level=…)/ZstdDecompressorlz4.frame.compress/lz4.frame.decompress
Wall time is time.perf_counter around the full buffer op (not streaming chunk APIs). MB/s = 64 MiB / p50_seconds. Ratio = raw_bytes / compressed_bytes. Round-trip equality was asserted every repeat.
text_real was built by walking /usr/share/doc, /usr/share, and /usr/lib/python3.13, skipping ELF/gzip, deduping by a content hash prefix, concatenating until ≥64 MiB, then truncating. That is still “lab box text,” not your production JSON — but it is far closer than a single repeated sentence.
If you re-run on NVMe with the zstd CLI -T0, expect higher absolute MB/s; keep the relative ordering as the portable lesson unless you re-measure.
Related links:
Verdict
On 64 MiB realistic /usr text, zstd level 3 delivered ratio 2.39 at ~227 MB/s compress and ~639 MB/s decompress — roughly 13× faster compress than gzip9 (2.02 / 17.8 / 169) with a better ratio. lz4 hit ~421 MB/s compress at ratio 1.58. urandom stayed at ≈1.000 ratio for all codecs. Choose the codec for the entropy you actually ship, and keep fairy-tale repeating buffers in the appendix.
Evidence path on the lab box: lab-evidence/33-zstd-vs-gzip-lz4/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5; python-zstandard 0.23.0 (CLI zstd 1.5.7); lz4 4.4.0+dfsg (CLI 1.10.0); gzip 1.13. Payloads 64 MiB. Lead text_real: gzip9 ratio 2.02 c 17.8 MB/s d 169; gzip6 2.01 / 28.4 / 164; zstd3 2.39 / 227 / 639; zstd1 2.00 / 365 / 624; zstd19 2.77 / 2.5 / 507; lz4 1.58 / 421 / 691. urandom ratio ≈1.000 all; zstd3 ~1181 MB/s compress. Pathological text_rep labeled. No Docker. Affiliates: 0. Evidence: lab-evidence/33-zstd-vs-gzip-lz4/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026