Plate 58
lzma vs bz2 Compress: Localhost Lab
Aditya Challa4 min read
Intro — what this post promises
Compress and decompress with lzma vs bz2, with zlib as a familiar reference (lab 102 covers zlib vs gzip). This lab reports MB/s and size ratios on Linux localhost.
Related links:
- reduce vs loop localhost lab
- selectors vs select localhost lab
- literal eval vs json localhost lab
- cmath vs math hypot localhost lab
- signal vs event wakeup localhost lab
- uuid4 vs uuid1 localhost lab
- platform vs uname localhost lab
- memoryview vs bytes localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Payload: 1048536 bytes of a repeated log line (highly compressible — ratios look extreme on purpose).
Verdict up front: compress — lzma ~62.48 MB/s (out 356 B); bz2 ~4.53 MB/s (out 550 B); zlib ref ~320.27 MB/s. Decompress: lzma ~1708.96, bz2 ~144.08.
Arms
| Arm | Pattern |
|---|---|
lzma.compress/decompress | XZ/LZMA2 defaults |
bz2.compress/decompress | bzip2 defaults |
zlib.compress(...,6) | reference only |
Seven rounds, p50. Round-trips OK (lzma=True, bz2=True).
Lab topology
Script: lab-evidence/143-lzma-vs-bz2/results/run_lab.py.
Lead table
| Arm | MB/s | out/in bytes | ratio |
|---|---|---|---|
| lzma compress | 62.48 | 356 | 0.0003 |
| bz2 compress | 4.53 | 550 | 0.0005 |
| zlib compress (ref) | 320.27 | 3656 | 0.0035 |
| lzma decompress | 1708.96 | 356 | — |
| bz2 decompress | 144.08 | 550 | — |
| zlib decompress (ref) | 1935.2 | 3656 | — |
Reading it for SRE work
- Archival / smallest artifacts on repetitive logs → lzma won both speed and size here.
- bz2 lagged badly on compress (~4.53 MB/s) — avoid for hot paths.
- Hot HTTP-ish payloads → stay on zlib/gzip (lab 102); this post is stdlib heavy compressors.
- Always measure on your corpus — random binary will not shrink like this repeated line.
Why ratios look tiny
The fixture repeats one log line to ~1 MiB. That is a realistic “agent dumps the same template” case, not a claim about JPEG. lzma’s 356-byte blob vs bz2’s 550 still shows lzma packing tighter on this pattern while compressing faster.
Decompress path
lzma decompress hit ~1708.96 MB/s vs bz2 ~144.08. Read-heavy pipelines that already store .xz stay fine; do not pick bz2 hoping decode will save you if encode was the bottleneck.
Put codec choice in the runbook next to retention: “xz for cold archive, gzip for hot ship.”
zlib reference context
zlib’s ~320.27 MB/s compress on this fixture is the speed ceiling among the three, with a larger 3656-byte blob. That matches the usual trade: gzip/zlib for hot ship, lzma for cold archive. Lab 102 remains the zlib-vs-gzip deep dive; this post only borrows zlib as a yardstick while crowning lzma over bz2 in the stdlib heavy tier.
On less repetitive data, absolute MB/s and ratios will move — rerun the script on a production sample before changing default codecs in an agent.
Pitfalls
- Comparing these MB/s to nginx gzip or zstd labs without stating codec.
- Using default presets without noting CPU vs size knobs.
- Compressing already-compressed blobs.
- Ignoring memory use of lzma presets on tiny hosts.
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box, default presets, one synthetic repetitive corpus. Not zstd (lab 33).
Takeaway
On this repetitive ~1 MiB log fixture, lzma compressed at ~62.48 MB/s to 356 bytes vs bz2 ~4.53 MB/s / 550 bytes. Prefer lzma over bz2 in stdlib; keep zlib/gzip for hotter paths.
Lab evidence
What I found running this
Ran the included lzma-versus-bz2 compression benchmark on Linux localhost with Python 3.13.5. Seven rounds used a 1,048,536-byte repeated log-line payload, measuring p50 compress and decompress throughput plus output-size ratios. lzma compressed faster and smaller than bz2 on this corpus, while zlib was the speed reference. Both codecs round-tripped successfully; random or already-compressed data could differ.
Related links
Plate 34
ast.literal_eval vs json.loads: Localhost Lab
1 Oct 2026
Plate 90
Path.read_text vs open().read: Localhost Lab
1 Oct 2026
Plate 68
gc.collect Cost Empty vs Cycles: Localhost Lab
Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.
Observability & SRE · 1 Oct 2026