ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 33

  1. Blog
  2. /Observability & SRE

SHA-256 vs BLAKE2b vs xxHash: Hash Throughput Lab

Hands-on SHA-256 vs BLAKE2b vs xxHash hash throughput on 64MiB buffers (SHA-1/SHA3 labeled). Real MB/s with OpenSSL and xxHash versions on the lab box.

Aditya Challa·30 September 2026·6 min read

Lab
On this page
  1. Intro — what this post promises
  2. What we compare (and what we do not)
  3. Lab topology
  4. Lead table — one-shot 64 MiB (p50)
  5. Streaming (1 MiB updates)
  6. How to read these numbers
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Versions pinned for this run
  10. Methodology footnote
  11. Relative speed (same 64 MiB buffer)
  12. Verdict

Intro — what this post promises

“Use BLAKE2, it’s faster than SHA-256” is a slide that aged poorly on machines with SHA-NI. Non-crypto checksums like xxHash live in a different league entirely — as long as you do not pretend they are MACs.

This is a hands-on lab with measured numbers:

  1. One-shot digest throughput over a 64 MiB os.urandom buffer.
  2. SHA-256, BLAKE2b (64- and 32-byte digests), BLAKE2s, SHA-1 (speed only), SHA3-256.
  3. xxHash family: xxh64, xxh128, xxh3_64, xxh3_128.
  4. Streaming with 1 MiB update() chunks for sha256 / blake2b / xxh64.

Related links:

  • zstd vs gzip vs lz4 compression localhost lab
  • openssl TLS full vs resume localhost lab
  • fsync vs fdatasync localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • process vs thread pool GIL localhost lab

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12, CPU flags include sha_ni). Python 3.13.5; OpenSSL 3.5.7; xxHash 0.8.3 via python3-xxhash. No Docker. Affiliates: 0. Throughput microbench — not a collision study, not HMAC.

Verdict up front: SHA-256 ~1468 MB/s beat BLAKE2b ~651 MB/s (~2.3×) on this box. xxh128 ~6695 MB/s led the checksum pack; xxh3_64 ~5998 MB/s.


What we compare (and what we do not)

AlgoRole in this postDigest size here
SHA-256crypto hash (OpenSSL)32 B
BLAKE2bcrypto hash (hashlib)64 B default + 32 B arm
BLAKE2scrypto hash32 B
SHA-1speed contrast only20 B
SHA3-256crypto hash32 B
xxHash / xxh3non-crypto checksum8–16 B

Related links:

  • openssl TLS full vs resume localhost lab

Lab topology

64 MiB os.urandom in RAM
One-shot: hashlib.<algo>(buf).digest() or xxhash.xxh*(buf).digest()
Stream: 1 MiB h.update slices, then digest()
Repeats: 7 (stream arms 5) after warmup; MB/s = 64 MiB / p50_s

Script: lab-evidence/35-sha256-vs-blake2-xxhash/results/run_lab.py.


Lead table — one-shot 64 MiB (p50)

AlgoMB/s p50p50 wallDigest
sha2561467.743.6 ms32
blake2b_64650.798.4 ms64
blake2b_32637.7100.4 ms32
blake2s_32423.6151.1 ms32
sha1 (speed only)1600.440.0 ms20
sha3_256429.0149.2 ms32
xxh644595.313.9 ms8
xxh1286694.79.6 ms16
xxh3_645998.210.7 ms8
xxh3_1285718.811.2 ms16

SHA-256 / BLAKE2b ≈ 2.25× on this sha_ni host. That inverts older “BLAKE2 always wins on software” intuition — measure your CPU.

xxHash sits ~3–4.5× above SHA-256 here. Fine for hash maps / content-defined chunking when cryptography is not the requirement.

Stability×3 MB/s p50: sha256 1462 / 1477 / 1475; blake2b_64 654 / 641 / 644; xxh3_64 5756 / 6271 / 6472 (noisier — quote bands).

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Streaming (1 MiB updates)

AlgoMB/s p50
sha256_stream_1mib1175.6
blake2b_stream_1mib609.3
xxh64_stream_1mib4969.5

Chunked update() costs a bit vs one-shot for SHA-256 (~20% here) and stays in the same ballpark for xxh64. Streaming is what real file/hash pipelines do — still SHA-256 ≫ BLAKE2b on this box.


How to read these numbers

  • Hardware accel matters. sha_ni in /proc/cpuinfo is the smoking gun for SHA-256’s lead.
  • Digest size ≠ speed. blake2b_32 did not beat blake2b_64 meaningfully here.
  • xxHash ≠ SHA-256. Different threat model; different MB/s class.
  • SHA-1 looked fast (~1600 MB/s) — still do not use it for security.
  • SHA3-256 (~429 MB/s) was the slow crypto hash in this set.

Related links:

  • zstd vs gzip vs lz4 compression localhost lab
  • nginx gzip on vs off localhost lab
  • asyncio vs threads IO concurrency lab

Pitfalls we hit (or avoided)

  1. Assuming BLAKE2b > SHA-256 everywhere — false on this OpenSSL + SHA-NI box.
  2. Calling xxHash a “faster SHA-256” — category error.
  3. Omitting library versions — OpenSSL 3.5.7, xxHash 0.8.3, Python 3.13.5.
  4. Tiny buffers — 64 MiB keeps fixed overhead small; still not a disk-bound hash.
  5. Multi-thread CLI tools — this is single-buffer Python; openssl dgst -sha256 multi-buffer may differ.

One more honesty note: these MB/s figures are single-threaded Python calling into C/OpenSSL/xxHash. A Rust or C xxh3 loop, or openssl speed sha256, can print different absolute numbers. Keep the ranking and the sha_ni caveat as the portable lessons unless you re-bench the exact binary you ship.

Practical checklist

  • For integrity/crypto: prefer SHA-256 or BLAKE2/SHA-3 with an explicit reason — benchmark the deploy CPU.
  • For non-crypto speed (hash tables, dedup fingerprints): consider xxh3 / xxHash — and document “not a MAC.”
  • Record OpenSSL / xxHash / CPU flags next to any MB/s claim.
  • Prefer streaming APIs for large files; spot-check vs one-shot.
  • Keep SHA-1 out of new security designs even when it wins a race.

Related links:

  • cosign SBOM CI sign/attest workflow
  • after xz supply-chain checklist
  • context-switch pipe ping-pong localhost lab

Versions pinned for this run

  • Python 3.13.5
  • OpenSSL 3.5.7 (9 Jun 2026)
  • xxHash library 0.8.3 (python3-xxhash 3.2.0-1+b5)
  • CPU advertises sha_ni among other AVX-512 / AES flags

If your OpenSSL build is older or your CPU lacks SHA extensions, expect SHA-256 to drop relative to BLAKE2b — that is a feature of the hardware, not a reason to distrust this lab’s method.

Methodology footnote

All crypto digests went through Python hashlib (OpenSSL backend for SHA-*). xxHash used the xxhash module’s xxh64 / xxh128 / xxh3_* constructors. Wall time is time.perf_counter around the full digest call. Input was os.urandom so content compressibility is irrelevant (unlike the compression lab).



Relative speed (same 64 MiB buffer)

Using SHA-256 p50 as 1.0×:

Algo≈ × vs SHA-256
sha1 (speed only)1.09×
sha2561.00×
blake2b_640.44×
sha3_2560.29×
xxh643.13×
xxh3_644.09×
xxh1284.56×

If your fleet is mixed (some VMs without SHA-NI, some with), do not copy this table blindly — re-run the same script on a canary of each SKU. OpenSSL version bumps can also move SHA-256 without touching your app code.

Related links:

  • SO_REUSEPORT vs single listen lab
  • tcp nodelay vs nagle localhost lab

Verdict

On this sha_ni box, SHA-256 (~1468 MB/s) outran BLAKE2b (~651 MB/s) by ~2.3×, while xxh128 (~6695 MB/s) and xxh3_64 (~5998 MB/s) owned non-crypto checksum throughput. Quote CPU flags and library versions with any hash MB/s number — folklore does not ship.

Evidence path on the lab box: lab-evidence/35-sha256-vs-blake2-xxhash/results/. Affiliates: 0.

sha-256 throughputblake2b vs sha256xxhash benchmarkxxh3hashlib pythonlocalhost labsrechecksum performance

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; OpenSSL 3.5.7; xxHash 0.8.3 (python3-xxhash). 64 MiB os.urandom, one-shot digest p50: sha256 1468 MB/s; blake2b-64 651; blake2b-32 638; blake2s 424; sha1 1600 (speed only); sha3_256 429; xxh64 4595; xxh3_64 5998; xxh128 6695; xxh3_128 5719. Stream 1 MiB chunks: sha256 1176; blake2b 609; xxh64 4970. CPU has sha_ni — SHA-256 beat BLAKE2b here (~2.3x). No Docker. Affiliates: 0. Evidence: lab-evidence/35-sha256-vs-blake2-xxhash/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compare (and what we do not)
  3. Lab topology
  4. Lead table — one-shot 64 MiB (p50)
  5. Streaming (1 MiB updates)
  6. How to read these numbers
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Versions pinned for this run
  10. Methodology footnote
  11. Relative speed (same 64 MiB buffer)
  12. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove