Plate 33
SHA-256 vs BLAKE2b vs xxHash: Hash Throughput Lab
Hands-on SHA-256 vs BLAKE2b vs xxHash hash throughput on 64MiB buffers (SHA-1/SHA3 labeled). Real MB/s with OpenSSL and xxHash versions on the lab box.
Aditya Challa6 min read
On this page
- Intro — what this post promises
- What we compare (and what we do not)
- Lab topology
- Lead table — one-shot 64 MiB (p50)
- Streaming (1 MiB updates)
- How to read these numbers
- Pitfalls we hit (or avoided)
- Practical checklist
- Versions pinned for this run
- Methodology footnote
- Relative speed (same 64 MiB buffer)
- Verdict
Intro — what this post promises
“Use BLAKE2, it’s faster than SHA-256” is a slide that aged poorly on machines with SHA-NI. Non-crypto checksums like xxHash live in a different league entirely — as long as you do not pretend they are MACs.
This is a hands-on lab with measured numbers:
- One-shot digest throughput over a 64 MiB
os.urandombuffer. - SHA-256, BLAKE2b (64- and 32-byte digests), BLAKE2s, SHA-1 (speed only), SHA3-256.
- xxHash family: xxh64, xxh128, xxh3_64, xxh3_128.
- Streaming with 1 MiB
update()chunks for sha256 / blake2b / xxh64.
Related links:
- zstd vs gzip vs lz4 compression localhost lab
- openssl TLS full vs resume localhost lab
- fsync vs fdatasync localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
- process vs thread pool GIL localhost lab
Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12, CPU flags include sha_ni). Python 3.13.5; OpenSSL 3.5.7; xxHash 0.8.3 via python3-xxhash. No Docker. Affiliates: 0. Throughput microbench — not a collision study, not HMAC.
Verdict up front: SHA-256 ~1468 MB/s beat BLAKE2b ~651 MB/s (~2.3×) on this box. xxh128 ~6695 MB/s led the checksum pack; xxh3_64 ~5998 MB/s.
What we compare (and what we do not)
| Algo | Role in this post | Digest size here |
|---|---|---|
| SHA-256 | crypto hash (OpenSSL) | 32 B |
| BLAKE2b | crypto hash (hashlib) | 64 B default + 32 B arm |
| BLAKE2s | crypto hash | 32 B |
| SHA-1 | speed contrast only | 20 B |
| SHA3-256 | crypto hash | 32 B |
| xxHash / xxh3 | non-crypto checksum | 8–16 B |
Related links:
Lab topology
Script: lab-evidence/35-sha256-vs-blake2-xxhash/results/run_lab.py.
Lead table — one-shot 64 MiB (p50)
| Algo | MB/s p50 | p50 wall | Digest |
|---|---|---|---|
| sha256 | 1467.7 | 43.6 ms | 32 |
| blake2b_64 | 650.7 | 98.4 ms | 64 |
| blake2b_32 | 637.7 | 100.4 ms | 32 |
| blake2s_32 | 423.6 | 151.1 ms | 32 |
| sha1 (speed only) | 1600.4 | 40.0 ms | 20 |
| sha3_256 | 429.0 | 149.2 ms | 32 |
| xxh64 | 4595.3 | 13.9 ms | 8 |
| xxh128 | 6694.7 | 9.6 ms | 16 |
| xxh3_64 | 5998.2 | 10.7 ms | 8 |
| xxh3_128 | 5718.8 | 11.2 ms | 16 |
SHA-256 / BLAKE2b ≈ 2.25× on this sha_ni host. That inverts older “BLAKE2 always wins on software” intuition — measure your CPU.
xxHash sits ~3–4.5× above SHA-256 here. Fine for hash maps / content-defined chunking when cryptography is not the requirement.
Stability×3 MB/s p50: sha256 1462 / 1477 / 1475; blake2b_64 654 / 641 / 644; xxh3_64 5756 / 6271 / 6472 (noisier — quote bands).
Related links:
Streaming (1 MiB updates)
| Algo | MB/s p50 |
|---|---|
| sha256_stream_1mib | 1175.6 |
| blake2b_stream_1mib | 609.3 |
| xxh64_stream_1mib | 4969.5 |
Chunked update() costs a bit vs one-shot for SHA-256 (~20% here) and stays in the same ballpark for xxh64. Streaming is what real file/hash pipelines do — still SHA-256 ≫ BLAKE2b on this box.
How to read these numbers
- Hardware accel matters.
sha_niin/proc/cpuinfois the smoking gun for SHA-256’s lead. - Digest size ≠ speed. blake2b_32 did not beat blake2b_64 meaningfully here.
- xxHash ≠ SHA-256. Different threat model; different MB/s class.
- SHA-1 looked fast (~1600 MB/s) — still do not use it for security.
- SHA3-256 (~429 MB/s) was the slow crypto hash in this set.
Related links:
- zstd vs gzip vs lz4 compression localhost lab
- nginx gzip on vs off localhost lab
- asyncio vs threads IO concurrency lab
Pitfalls we hit (or avoided)
- Assuming BLAKE2b > SHA-256 everywhere — false on this OpenSSL + SHA-NI box.
- Calling xxHash a “faster SHA-256” — category error.
- Omitting library versions — OpenSSL 3.5.7, xxHash 0.8.3, Python 3.13.5.
- Tiny buffers — 64 MiB keeps fixed overhead small; still not a disk-bound hash.
- Multi-thread CLI tools — this is single-buffer Python;
openssl dgst -sha256multi-buffer may differ.
One more honesty note: these MB/s figures are single-threaded Python calling into C/OpenSSL/xxHash. A Rust or C xxh3 loop, or openssl speed sha256, can print different absolute numbers. Keep the ranking and the sha_ni caveat as the portable lessons unless you re-bench the exact binary you ship.
Practical checklist
- For integrity/crypto: prefer SHA-256 or BLAKE2/SHA-3 with an explicit reason — benchmark the deploy CPU.
- For non-crypto speed (hash tables, dedup fingerprints): consider xxh3 / xxHash — and document “not a MAC.”
- Record OpenSSL / xxHash / CPU flags next to any MB/s claim.
- Prefer streaming APIs for large files; spot-check vs one-shot.
- Keep SHA-1 out of new security designs even when it wins a race.
Related links:
- cosign SBOM CI sign/attest workflow
- after xz supply-chain checklist
- context-switch pipe ping-pong localhost lab
Versions pinned for this run
- Python 3.13.5
- OpenSSL 3.5.7 (9 Jun 2026)
- xxHash library 0.8.3 (
python3-xxhash3.2.0-1+b5) - CPU advertises sha_ni among other AVX-512 / AES flags
If your OpenSSL build is older or your CPU lacks SHA extensions, expect SHA-256 to drop relative to BLAKE2b — that is a feature of the hardware, not a reason to distrust this lab’s method.
Methodology footnote
All crypto digests went through Python hashlib (OpenSSL backend for SHA-*). xxHash used the xxhash module’s xxh64 / xxh128 / xxh3_* constructors. Wall time is time.perf_counter around the full digest call. Input was os.urandom so content compressibility is irrelevant (unlike the compression lab).
Relative speed (same 64 MiB buffer)
Using SHA-256 p50 as 1.0×:
| Algo | ≈ × vs SHA-256 |
|---|---|
| sha1 (speed only) | 1.09× |
| sha256 | 1.00× |
| blake2b_64 | 0.44× |
| sha3_256 | 0.29× |
| xxh64 | 3.13× |
| xxh3_64 | 4.09× |
| xxh128 | 4.56× |
If your fleet is mixed (some VMs without SHA-NI, some with), do not copy this table blindly — re-run the same script on a canary of each SKU. OpenSSL version bumps can also move SHA-256 without touching your app code.
Related links:
Verdict
On this sha_ni box, SHA-256 (~1468 MB/s) outran BLAKE2b (~651 MB/s) by ~2.3×, while xxh128 (~6695 MB/s) and xxh3_64 (~5998 MB/s) owned non-crypto checksum throughput. Quote CPU flags and library versions with any hash MB/s number — folklore does not ship.
Evidence path on the lab box: lab-evidence/35-sha256-vs-blake2-xxhash/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5; OpenSSL 3.5.7; xxHash 0.8.3 (python3-xxhash). 64 MiB os.urandom, one-shot digest p50: sha256 1468 MB/s; blake2b-64 651; blake2b-32 638; blake2s 424; sha1 1600 (speed only); sha3_256 429; xxh64 4595; xxh3_64 5998; xxh128 6695; xxh3_128 5719. Stream 1 MiB chunks: sha256 1176; blake2b 609; xxh64 4970. CPU has sha_ni — SHA-256 beat BLAKE2b here (~2.3x). No Docker. Affiliates: 0. Evidence: lab-evidence/35-sha256-vs-blake2-xxhash/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026