ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 76

  1. Blog

html.parser vs regex Tag Strip: Localhost Lab

Aditya Challa·1 October 2026·4 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 docs/s)
  5. Behavior differences
  6. Reading it for SRE work
  7. Throughput vs fidelity
  8. Scale note
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

Strip tags from HTML fragments with html.parser.HTMLParser vs regex. This lab reports docs/s on Linux localhost, plus script/entity behavior checks.

Related links:

  • copy copy vs dict copy localhost lab
  • itertools batched vs chunk localhost lab
  • exitstack vs nested with localhost lab
  • chainmap vs dict merge localhost lab
  • ipaddress vs string prefix localhost lab
  • fractions vs float localhost lab
  • math fsum vs sum localhost lab
  • xml etree vs json localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Differentiates from html-escape-vs-manual (lab 105) — that post escaped text; this one strips markup.

Verdict up front (small doc 324 bytes, n=5000): regex ~243373 docs/s; html.parser ~14389. Bigger doc (2600 B): regex ~36491 vs parser ~1866.


Arms

ArmPattern
HTMLParser subclassskip script/style; join data
regex stripdrop script/style blocks then tags
both on 8× docscale check

Seven rounds, p50. Fragment includes attrs, comment, entities, nested tags.


Lab topology

n_small=5000 · n_big=1250 · 7 rounds · p50
metric: docs/s = n / p50_s

Script: lab-evidence/127-html-parser-vs-regex/results/run_lab.py.


Lead table (p50 docs/s)

Armdocs/s
regex small243373
html.parser small14389
regex big36491
html.parser big1866

Regex led throughput by a wide margin on both sizes. Correctness is the other axis.


Behavior differences

  • Both arms dropped script bodies (alert not in text).
  • Parser decoded entities (& → &); regex left & / < in the stripped string.
  • Parser text (repr): '\n\nAlert severity\nLatency was 12.5 ms & rising — see link.\n\nonetwo <three>\n'
  • Regex text (repr): '\n\nAlert severity\nLatency was 12.5 ms &amp; rising — see link.\n\nonetwo &lt;three&gt;\n'

Fast regex is not a full HTML stack — nested oddities, malformed tags, and CDATA will diverge.


Reading it for SRE work

  • Best-effort plain text from trusted snippets → regex can be enough if you accept entity leftovers.
  • Need entity decode + structured skip of script/style → html.parser (or a dedicated library you measure later).
  • Never use either as an XSS sanitizer for untrusted HTML in browsers — wrong tool.
  • Lab 105 covers escaping for safe display; this post is strip-to-text only.

Throughput vs fidelity

On the small fragment, regex ran about 16.9× the parser docs/s (~243373 vs ~14389). That gap buys you speed, not a DOM. If on-call pastes alert HTML into a ticket field, prefer the parser when humans must read decoded &amp; as &.

Pair strip choice with an explicit unit fixture for script leak and entity decode — speed alone will pick regex every time and hide the fidelity bug.



Scale note

On the 8× document (2600 bytes), regex still led (~36491 docs/s) while the parser sat near ~1866. Absolute docs/s fell for both as expected; the fidelity gap (entities) did not close with size. Measure your own corpus before codifying a strip helper in a shared library.


Pitfalls

  • Tag regex breaking on > inside attributes.
  • Forgetting script/style pre-strip and leaking JS into “text”.
  • Treating strip output as safe HTML.
  • Comparing to lab 105 escape numbers (different problem).

Reproduce

python3 lab-evidence/127-html-parser-vs-regex/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box. Stdlib only. Not BeautifulSoup/lxml, not XSS filters.


Takeaway

Regex strip hit ~243373 docs/s vs parser ~14389 on the small doc — but the parser decoded entities. Pick regex for speed on simple trusted markup; pick html.parser when decode and tag structure matter.

pythonhtml.parserregexbenchmarkingperformancehtmlparsing

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. small: regex 243373 docs/s; html.parser 14389. Parser decodes entities; regex left &. Affiliates: 0. Evidence: lab-evidence/127-html-parser-vs-regex/. script_leaked=False on both inputs.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 34

    ast.literal_eval vs json.loads: Localhost Lab

    1 Oct 2026

  • Plate 68

    gc.collect Cost Empty vs Cycles: Localhost Lab

    Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 77

    fractions.Fraction vs float: Localhost Lab

    Hands-on fractions.Fraction vs float for exact ratios: real ops/s and exactness checks, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table (p50 docs/s)
  5. Behavior differences
  6. Reading it for SRE work
  7. Throughput vs fidelity
  8. Scale note
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove