ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 58

  1. Blog
  2. /Observability & SRE

Python re vs str Methods: Text Scan Throughput Lab

Hands-on Python re vs str methods lab on a 32MiB localhost corpus: count, find, line prefix, case-insensitive, and email scans with real measured MB/s.

Aditya Challa·30 September 2026·6 min read

Lab
On this page
  1. Intro — what this post promises
  2. Lab topology
  3. Arm A — literal count (`shopper`)
  4. Arm B — find / search (do not trust early hits)
  5. Arm C — line prefix `ERROR`
  6. Arm D — case-insensitive + structured extracts
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Methodology footnote
  11. Ratio cheat sheet (time\_re / time\_str; >1 means re slower)
  12. Versions / environment pinned
  13. Verdict

Intro — what this post promises

Reach for re when you need patterns. Reach for str when you need speed on literals. This lab puts numbers on that slogan over a ~32 MiB synthetic log-ish corpus.

This is a hands-on lab with measured numbers:

  1. Literal count — str.count vs re.findall / finditer.
  2. find / search — including missing and late needles (full scans).
  3. Line prefix — startswith scan vs re.compile(r'^ERROR', re.M).
  4. Case-insensitive and “email-ish” / capture-group shapes.

Related links:

  • json vs orjson vs msgpack localhost lab
  • SHA-256 vs BLAKE2b vs xxHash localhost lab
  • zstd vs gzip vs lz4 compression localhost lab
  • stdout buffering line vs full localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (1 Oct 2026 IST): Shared Linux lab box. Python 3.13.5. Corpus ~32 MiB, ~395k lines, written to evidence. Affiliates: 0. Microbench on synthetic text — not production PCRE2, not a parser.

Verdict up front: str.count ~973 MB/s vs re.findall ~637 MB/s on a literal (1.53×). Prefix ERROR lines: str scan ~1016 MB/s vs multiline re ~219 MB/s (4.6×). Simple capture extract nearly tied.


Lab topology

Synthetic INFO/ERROR/GET/email lines → ~32 MiB text
precompiled re patterns where used
bench: warmup + 7 timed calls; MB/s = corpus_bytes / p50_s
Assert matching counts between paired str/re arms

Script: lab-evidence/39-python-re-vs-str/results/run_lab.py.


Arm A — literal count (shopper)

Methodp50MB/shits
str.count32.9 ms973409861
re.findall50.2 ms637409861
re.finditer count59.0 ms542409861

re.findall / str.count time ≈ 1.53×. Same hit count — fair fight.

Related links:

  • Why your average latency graph is lying (p50 / p95 / p99)

Arm B — find / search (do not trust early hits)

First occurrence of shopper sits near the start — both APIs return in microseconds; MB/s looks absurd. Full-scan arms:

Methodp50MB/sresult
str.find missing10.5 ms3039-1
re.search missing22.3 ms1438none
str.find late needle7.0 ms4554EOF offset
re.search late21.2 ms1509same

Missing: re ~2.1× slower. Late: re ~3.0× slower. Use these, not the early-hit toy.


Arm C — line prefix ERROR

Methodp50MB/slines
splitlines + startswith65.1 ms4927899
manual scan + startswith31.5 ms10167899
re.findall(r'^ERROR', re.M)146 ms2197899

Avoid building a giant splitlines list when a scan suffices. Multiline ^ regex paid ~4.6× vs the scan.


Arm D — case-insensitive + structured extracts

Methodp50MB/s
text.lower().count('shopper')57.2 ms560
re.findall(..., re.I)337 ms94.9
naive str email walker39.3 ms814
re.findall email pattern636 ms50.3
str scan GET /api/v1/items/<id>20.8 ms1541
re.findall with group22.8 ms1403

re.I ~5.9× slower than lower+count here (and lower allocates a full copy — still won). Email regex is more expressive; the ~16× gap is “pattern power costs CPU.” GET id extract is where a tight regex nearly matches a hand scanner.

Related links:

  • json vs orjson vs msgpack localhost lab
  • SHA-256 vs BLAKE2b vs xxHash localhost lab
  • stdout buffering line vs full localhost lab

How to read these numbers

  • Literals → str. count/find/startswith are hard to beat.
  • Real patterns → re. Pay the tax knowingly (emails, alternation, lookarounds).
  • Compile once. All re arms used re.compile.
  • MB/s assumes a full corpus pass — early-exit finds need separate labeling.

Pitfalls we hit (or avoided)

  1. O(n²) corpus builder (sum(len(line)) each append) — hung the first run; fixed with a running counter.
  2. Citing early find MB/s as throughput — nonsense; added missing/late arms.
  3. Comparing splitlines + startswith to regex without a scan baseline.
  4. Calling naive email walker “equivalent” to RFC regex — same hit count here, different power.
  5. Forgetting re.I vs explicit .lower() allocation tradeoffs.

Practical checklist

  • Hot path literal scans: prefer str methods.
  • Need a grammar: use re, precompile, measure.
  • Prefer scanning over splitlines() on large blobs when possible.
  • For casefold counting, bench .lower() vs re.I on your data.
  • Keep paired asserts so str/re refactors do not drift on counts.

Related links:

  • process vs thread pool GIL localhost lab
  • asyncio vs threads IO concurrency lab
  • posix fadvise sequential vs random localhost lab
  • fork COW RSS vs spawn localhost lab

Methodology footnote

Corpus mixes INFO filler, periodic ERROR lines, GET paths, and userN@example.com lines. Timed with time.perf_counter. MB/s = UTF-8 byte length / p50 seconds. Result counts asserted equal across paired arms (count, ERROR lines, GET ids).


Ratio cheat sheet (time_re / time_str; >1 means re slower)

TaskRatio
literal findall vs count1.53×
missing search vs find2.11×
late search vs find3.02×
^ERROR vs scan startswith~4.6×
re.I findall vs lower+count~5.9×
email findall vs naive walker~16×
GET id findall vs scan~1.10×

When the ratio approaches 1.0, readability may win and you keep the regex. When it is 4×+ on a hot path, write the literal scanner (with tests) or narrow the regex.

For multi-gigabyte logs, consider streaming line-by-line (still str.startswith / compiled re on each line) so you do not hold a 32 MiB+ contiguous unicode object. Throughput will drop; peak RSS will thank you — see the fork COW lab for why large private buffers hurt.

Versions / environment pinned

  • Python 3.13.5 stdlib re (no third-party regex engine)
  • Corpus ~32 MiB / ~395k lines at lab-evidence/39-python-re-vs-str/results/corpus.txt
  • Patterns precompiled; benches use p50 of 7 timed calls after warmup

Bottom line for Python services: default to str for fixed tokens and reserved words in logs; promote to re when the token shape is genuinely a language. The GET-id near-parity case shows a careful regex need not be slow — it also shows you should still verify against a scanner before declaring victory.

Verdict

On a ~32 MiB corpus, literal str.count (~973 MB/s) beat re.findall (~637 MB/s) by ~1.53×, and an ERROR-prefix scan beat multiline regex by ~4.6×. A simple GET-id regex nearly matched a hand scanner. Use str for literals; keep re for patterns — and measure the boundary.

Evidence path on the lab box: lab-evidence/39-python-re-vs-str/results/. Affiliates: 0.

python re vs strstr.countre.findallregex performancestartswithlocalhost labsretext scanning

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5; ~32 MiB synthetic corpus (~395k lines). str.count('shopper') 973 MB/s vs re.findall 637 (~1.53x slower). ERROR line prefix: str scan 1016 MB/s vs re^ERROR 219 (~4.6x). Case-insensitive: str.lower().count 560 vs re.I 95 (~5.9x). Missing needle full scan: str.find 3039 MB/s vs re.search 1438 (~2.1x). Emails: naive str 814 vs re 50 (~16x; pattern power differs). GET id extract: str 1541 vs re 1403 (close). Affiliates: 0. Evidence: lab-evidence/39-python-re-vs-str/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Lab topology
  3. Arm A — literal count (`shopper`)
  4. Arm B — find / search (do not trust early hits)
  5. Arm C — line prefix `ERROR`
  6. Arm D — case-insensitive + structured extracts
  7. How to read these numbers
  8. Pitfalls we hit (or avoided)
  9. Practical checklist
  10. Methodology footnote
  11. Ratio cheat sheet (time\_re / time\_str; >1 means re slower)
  12. Versions / environment pinned
  13. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove