ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 56

  1. Blog

json vs orjson vs msgpack: Python Serializer Lab

Hands-on Python json vs orjson vs msgpack lab: encode/decode MB/s on records, nested, and unicode payloads. Real p50 numbers measured on localhost box.

Aditya Challa·30 September 2026·6 min read

Summary
On this page
  1. Intro — what this post promises
  2. What each codec is
  3. Lab topology
  4. Lead table — records\_5k (rebench n=25, p50)
  5. Nested config + mixed unicode (first-pass matrix)
  6. How to read these numbers
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Versions pinned for this run
  10. Encode RPS view (records\_5k)
  11. Methodology footnote
  12. When stdlib json is still fine
  13. Verdict

Intro — what this post promises

Python’s stdlib json is everywhere. orjson and msgpack promise faster encode/decode and (for msgpack) smaller blobs. How much on realistic dict/list shapes?

This is a hands-on lab with measured numbers:

  1. Encode + decode throughput for json, orjson, and msgpack.
  2. Three payloads: 5 000 API-ish records, nested config, mixed unicode document.
  3. Size on the wire (bytes) next to MB/s.
  4. A denser json compact arm so orjson is not winning only on whitespace.

Related links:

  • SHA-256 vs BLAKE2b vs xxHash localhost lab
  • zstd vs gzip vs lz4 compression localhost lab
  • SQLite WAL vs DELETE journal localhost lab
  • fsync vs fdatasync localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5; orjson 3.10.7; msgpack 1.0.3. No Docker. Affiliates: 0. In-process microbench — not HTTP framing, not schema evolution.

Verdict up front: on records_5k, orjson encoded at ~542 MB/s vs stdlib json ~100 MB/s (~5.4×). msgpack was the smallest (285 KiB vs json 453 KiB) at ~181 MB/s encode.


What each codec is

CodecFormatNotes in this lab
jsontext JSONjson.dumps(...).encode(); also compact separators
orjsontext JSONreturns bytes; default compact
msgpackbinary MessagePackuse_bin_type=True, raw=False on unpack

Related links:

  • zstd vs gzip vs lz4 compression localhost lab

Lab topology

Payloads in RAM → encode → decode → light structural assert
records_5k     — 5000 dict rows (id, sku, price, tags, active)
nested_config  — 200 services with env + peers + small matrix
mixed_unicode  — unicode title/body + 2000 short items
Metric: p50 wall → MB/s = encoded_bytes / p50_s
Lead table: records_5k rebench n=25 after warmup

Script: lab-evidence/36-json-vs-orjson-msgpack/results/run_lab.py.


Lead table — records_5k (rebench n=25, p50)

CodecBytesEncode MB/sDecode MB/sEncode p50
json (default)452 96299.8121.34.33 ms
json compact396 96388.4109.34.28 ms
orjson396 963542.3209.60.70 ms
msgpack284 619180.6109.31.50 ms

orjson / json encode ≈ 5.4×. Decode ≈ 1.7×. Compact stdlib matched orjson’s byte size but stayed in the ~90 MB/s encode band — the win is the library, not missing spaces.

msgpack shrank ~37% vs default json and ~28% vs compact JSON, with encode ~1.8× json.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Nested config + mixed unicode (first-pass matrix)

Payload / codecBytesEnc MB/sDec MB/s
nested / json3346175.580.2
nested / orjson28856445.0134.6
nested / msgpack1911998.651.8
mixed / json10986891.3129.1
mixed / orjson99461749.2202.2
mixed / msgpack83054248.9192.2

orjson’s encode lead held across shapes (nested ~5.9×, mixed ~8.2× vs json on that pass). msgpack stayed the size champion.


How to read these numbers

  • Encode-bound APIs (building responses): orjson is the blunt instrument here.
  • Size-bound links (queues, disks): msgpack earned the byte win; pair with the compression lab if you also gzip/zstd.
  • Decode gaps were smaller than encode on records — measure both directions.
  • stdlib compact ≠ orjson speed.

Related links:

  • SQLite WAL vs DELETE journal localhost lab
  • fsync vs fdatasync localhost lab
  • SHA-256 vs BLAKE2b vs xxHash localhost lab

Pitfalls we hit (or avoided)

  1. Comparing pretty-printed json to orjson — we also ran compact separators.
  2. One noisy pass — early orjson encode on records looked ~306 MB/s; n=25 rebench settled ~542 MB/s. Quote the rebench.
  3. Assuming msgpack always faster than json — size yes; encode was between json and orjson here.
  4. HTTP Content-Type lies — binary msgpack needs clients that speak it.
  5. Type fidelity — msgpack/orjson edge cases (tuple keys, datetime) not the focus.

Practical checklist

  • If you emit lots of JSON from Python, trial orjson on a production-shaped fixture.
  • If bytes dominate cost, trial msgpack (or JSON + zstd — see compression lab).
  • Benchmark encode and decode with p50/p95, not one timeit line.
  • Pin versions (orjson 3.10.7, msgpack 1.0.3 here).
  • Keep evidence next to any “we switched serializers” capacity claim.

Group commit for serializers is “batch your documents”: amortize fixed costs the same way the SQLite lab amortizes commits.

Related links:

  • nginx gzip on vs off localhost lab
  • process vs thread pool GIL localhost lab
  • asyncio vs threads IO concurrency lab

Versions pinned for this run

  • Python 3.13.5
  • orjson 3.10.7 (python3-orjson)
  • msgpack 1.0.3 (python3-msgpack)


Encode RPS view (records_5k)

MB/s divides by encoded size; RPS asks “how many full document dumps per second?”

Using rebench p50 encode times:

CodecEncode p50≈ docs/s
json4.33 ms~231
orjson0.70 ms~1429
msgpack1.50 ms~667

If your service builds one fat JSON per request, orjson’s ~6× docs/s vs json is the capacity story — before HTTP and disk enter the chat (see keepalive / fsync labs).

Stability encode MB/s on records (earlier 3×5 pass): json ~87–99, orjson ~506–528, msgpack ~164–178 — same ranking as the n=25 rebench.

Methodology footnote

Encoders:

  • json.dumps(obj).encode("utf-8") and compact separators=(",", ":")
  • orjson.dumps(obj) / orjson.loads
  • msgpack.packb(..., use_bin_type=True) / unpackb(..., raw=False)

Wall clocks use time.perf_counter. MB/s divides encoded byte length by p50 seconds (decode uses the same size denominator so encode/decode MB/s are comparable for a given codec). The n=25 rebench is what the lead table quotes after a noisier first pass.

If you already compress JSON with gzip/zstd at the edge, re-bench end-to-end: a smaller msgpack body can beat orjson+JSON on the wire even when orjson wins CPU encode. The compression and nginx-gzip labs are the right companions for that follow-up.

When stdlib json is still fine

Not every service is encode-bound. If you dump a 2 KiB dict once per request behind 50 ms of ORM time, orjson will not move p95. This lab is for the hot path: large fan-out JSON, analytics exports, cache fills, and chatty internal RPCs that serialize constantly. Profile first; then swap the codec and re-measure with the same fixtures we used (or better — your fixtures).

Verdict

On 5 000 record payloads, orjson encoded at ~542 MB/s versus stdlib json ~100 MB/s (~5.4×) and decoded ~1.7× faster. msgpack delivered the smallest blobs (~285 KiB vs ~453 KiB json) at ~181 MB/s encode. Pick orjson for JSON speed, msgpack for binary size — and measure your shape.

Evidence path on the lab box: lab-evidence/36-json-vs-orjson-msgpack/results/. Affiliates: 0.

orjson vs jsonmsgpack pythonjson encode throughputserializer benchmarkorjsonlocalhost labsrepython performance

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; orjson 3.10.7; msgpack 1.0.3. Lead records_5k rebench n=25: json enc 99.8 MB/s dec 121.3; orjson enc 542.3 (~5.4x) dec 209.6 (~1.7x); msgpack enc 180.6 dec 109.3 size 285 KiB (smallest) vs json 453 / orjson 397. nested orjson enc ~445 MB/s vs json ~76; mixed unicode orjson enc ~749. Affiliates: 0. Evidence: lab-evidence/36-json-vs-orjson-msgpack/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 78

    json.dumps Compact vs Indent Lab

    Hands-on json.dumps separators vs indent lab: real time and byte-size tradeoffs for compact vs pretty JSON encoding, measured on Linux localhost (lab).

    Observability & SRE · 30 Sept 2026

  • Plate 29

    base64 vs hex Encode/Decode Throughput Lab

    Hands-on base64 vs hex encode/decode lab on 32MiB buffers with real MB/s measurements in both directions on localhost.

    30 Sept 2026

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What each codec is
  3. Lab topology
  4. Lead table — records\_5k (rebench n=25, p50)
  5. Nested config + mixed unicode (first-pass matrix)
  6. How to read these numbers
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Versions pinned for this run
  10. Encode RPS view (records\_5k)
  11. Methodology footnote
  12. When stdlib json is still fine
  13. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove