ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 58

  1. Blog

lzma vs bz2 Compress: Localhost Lab

Aditya Challa·1 October 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table
  5. Reading it for SRE work
  6. Why ratios look tiny
  7. Decompress path
  8. zlib reference context
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

Compress and decompress with lzma vs bz2, with zlib as a familiar reference (lab 102 covers zlib vs gzip). This lab reports MB/s and size ratios on Linux localhost.

Related links:

  • reduce vs loop localhost lab
  • selectors vs select localhost lab
  • literal eval vs json localhost lab
  • cmath vs math hypot localhost lab
  • signal vs event wakeup localhost lab
  • uuid4 vs uuid1 localhost lab
  • platform vs uname localhost lab
  • memoryview vs bytes localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Payload: 1048536 bytes of a repeated log line (highly compressible — ratios look extreme on purpose).

Verdict up front: compress — lzma ~62.48 MB/s (out 356 B); bz2 ~4.53 MB/s (out 550 B); zlib ref ~320.27 MB/s. Decompress: lzma ~1708.96, bz2 ~144.08.


Arms

ArmPattern
lzma.compress/decompressXZ/LZMA2 defaults
bz2.compress/decompressbzip2 defaults
zlib.compress(...,6)reference only

Seven rounds, p50. Round-trips OK (lzma=True, bz2=True).


Lab topology

raw=1048536 B · 7 rounds · p50
MB/s = (raw_MiB) / p50_s

Script: lab-evidence/143-lzma-vs-bz2/results/run_lab.py.


Lead table

ArmMB/sout/in bytesratio
lzma compress62.483560.0003
bz2 compress4.535500.0005
zlib compress (ref)320.2736560.0035
lzma decompress1708.96356—
bz2 decompress144.08550—
zlib decompress (ref)1935.23656—

Reading it for SRE work

  • Archival / smallest artifacts on repetitive logs → lzma won both speed and size here.
  • bz2 lagged badly on compress (~4.53 MB/s) — avoid for hot paths.
  • Hot HTTP-ish payloads → stay on zlib/gzip (lab 102); this post is stdlib heavy compressors.
  • Always measure on your corpus — random binary will not shrink like this repeated line.

Why ratios look tiny

The fixture repeats one log line to ~1 MiB. That is a realistic “agent dumps the same template” case, not a claim about JPEG. lzma’s 356-byte blob vs bz2’s 550 still shows lzma packing tighter on this pattern while compressing faster.


Decompress path

lzma decompress hit ~1708.96 MB/s vs bz2 ~144.08. Read-heavy pipelines that already store .xz stay fine; do not pick bz2 hoping decode will save you if encode was the bottleneck.

Put codec choice in the runbook next to retention: “xz for cold archive, gzip for hot ship.”



zlib reference context

zlib’s ~320.27 MB/s compress on this fixture is the speed ceiling among the three, with a larger 3656-byte blob. That matches the usual trade: gzip/zlib for hot ship, lzma for cold archive. Lab 102 remains the zlib-vs-gzip deep dive; this post only borrows zlib as a yardstick while crowning lzma over bz2 in the stdlib heavy tier.

On less repetitive data, absolute MB/s and ratios will move — rerun the script on a production sample before changing default codecs in an agent.


Pitfalls

  • Comparing these MB/s to nginx gzip or zstd labs without stating codec.
  • Using default presets without noting CPU vs size knobs.
  • Compressing already-compressed blobs.
  • Ignoring memory use of lzma presets on tiny hosts.

Reproduce

python3 lab-evidence/143-lzma-vs-bz2/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box, default presets, one synthetic repetitive corpus. Not zstd (lab 33).


Takeaway

On this repetitive ~1 MiB log fixture, lzma compressed at ~62.48 MB/s to 356 bytes vs bz2 ~4.53 MB/s / 550 bytes. Prefer lzma over bz2 in stdlib; keep zlib/gzip for hotter paths.

pythoncompressionlzmabz2zlibbenchmarkinglinuxstdlib

Lab evidence

What I found running this

Ran the included lzma-versus-bz2 compression benchmark on Linux localhost with Python 3.13.5. Seven rounds used a 1,048,536-byte repeated log-line payload, measuring p50 compress and decompress throughput plus output-size ratios. lzma compressed faster and smaller than bz2 on this corpus, while zlib was the speed reference. Both codecs round-tripped successfully; random or already-compressed data could differ.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 34

    ast.literal_eval vs json.loads: Localhost Lab

    1 Oct 2026

  • Plate 90

    Path.read_text vs open().read: Localhost Lab

    1 Oct 2026

  • Plate 68

    gc.collect Cost Empty vs Cycles: Localhost Lab

    Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table
  5. Reading it for SRE work
  6. Why ratios look tiny
  7. Decompress path
  8. zlib reference context
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove