ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 15

  1. Blog

xml.etree vs json Nested Records: Localhost Lab

Aditya Challa·1 October 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 ops/s)
  5. Correctness check
  6. Why XML still shows up
  7. Reading it for SRE work
  8. Encode vs decode skew
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

Serialize and parse nested telemetry-like records with json vs xml.etree.ElementTree. This lab reports ops/s and payload bytes on Linux localhost — same schema, same record count.

Related links:

  • math fsum vs sum localhost lab
  • fractions vs float localhost lab
  • chainmap vs dict merge localhost lab
  • ipaddress vs string prefix localhost lab
  • statistics quantiles vs manual localhost lab
  • random choices vs sample localhost lab
  • itertools batched vs chunk localhost lab
  • exitstack vs nested with localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Differentiates from configparser-vs-json (lab 104) and tomllib-vs-json (lab 112) — here the peer is XML ElementTree, not INI/TOML.

Verdict up front (n=2000 records): json.dumps ~673887 ops/s; xml build+tostring ~158580; json.loads ~1520987; xml fromstring+parse ~349700. Bytes: JSON 172948, XML 248299.


Arms

ArmPattern
json.dumps (compact)encode list[dict]
ElementTree build + tostringencode equivalent XML
json.loadsdecode
ET.fromstring + walkdecode to dicts
round-trip eachencode+decode

Seven rounds, p50. Each record: id, host, latency_ms, ok, three tags.


Lab topology

n=2000 records · 7 rounds · p50
metric: ops/s = n / p50_s

Script: lab-evidence/124-xml-etree-vs-json/results/run_lab.py.


Lead table (p50 ops/s)

Armops/s
json.dumps673887
xml build+tostring158580
json.loads1520987
xml fromstring+parse349700
json round-trip676101
xml round-trip106991

JSON wins every throughput column on this box. XML also grew the wire: 248299 bytes vs JSON 172948 (~1.44×).


Correctness check

Both paths returned 2000 / 2000 records; first id and last host matched across formats. The XML arm is a fair structural peer, not a toy one-liner.


Why XML still shows up

Vendors, SOAP leftovers, and device configs still ship XML. The cost is real: round-trip XML sat at ~106991 ops/s vs JSON ~676101. Prefer JSON for new internal envelopes; keep ElementTree when the contract is XML and you cannot change it.


Reading it for SRE work

  • Internal service payloads → json (faster encode/decode, smaller blob here).
  • Must speak XML → ElementTree (or lxml if you later measure it); budget CPU and bytes.
  • Do not “translate to XML for logs” without a contract — you pay about 4.2× encode slowdown on this run (json ops/s ÷ xml ops/s).
  • Labs 104/112 cover INI/TOML peers; this post is XML vs JSON only.

Document the chosen envelope in the runbook so on-call does not “normalize everything to XML” under incident pressure.



Encode vs decode skew

On this run, decode was the friendlier XML column (~349700 ops/s) but still trailed json.loads (~1520987 ops/s). Encode hurt more: ElementTree element allocation plus tostring sat near ~158580 ops/s against json.dumps ~673887. If your pipeline is write-heavy (agent flush), XML tax shows up first; if it is read-heavy (config ingest), budget the parse arm and the larger 248299-byte blob on the wire.

Prefer measuring your schema — attribute-heavy XML can swing these ratios. This lab keeps a fixed five-field record so ops/s stay comparable across arms.


Pitfalls

  • Comparing pretty-printed JSON to minified XML (or the reverse).
  • Using regex to “parse” XML instead of ElementTree when namespaces appear.
  • Forgetting attribute vs text modeling differences when mapping to dicts.
  • Assuming third-party XML parsers match these ElementTree numbers.

Reproduce

python3 lab-evidence/124-xml-etree-vs-json/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box, stdlib only (no lxml). Synthetic records, not a vendor WSDL corpus.


Takeaway

On 2000 nested records, json led encode (~673887 ops/s) and decode (~1520987), with a smaller payload (172948 vs 248299 bytes). Use ElementTree when the wire format is XML; prefer JSON for new internal telemetry.

xml.etreeelementtreejsonserializepython xmllocalhost labsreops/s

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. n=2000: json.dumps 673887 ops/s; xml build 158580; json.loads 1520987; xml parse 349700. bytes json=172948 xml=248299. Affiliates: 0. Evidence: lab-evidence/124-xml-etree-vs-json/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 75

    uuid.uuid4 vs uuid.uuid1: Localhost Lab

    Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 76

    cmath vs math.hypot Magnitudes: Localhost Lab

    Hands-on cmath vs math.hypot magnitude ops lab: real ops/s for abs, polar, and phase, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 ops/s)
  5. Correctness check
  6. Why XML still shows up
  7. Reading it for SRE work
  8. Encode vs decode skew
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove