Plate 58
Python re vs str Methods: Text Scan Throughput Lab
Hands-on Python re vs str methods lab on a 32MiB localhost corpus: count, find, line prefix, case-insensitive, and email scans with real measured MB/s.
Aditya Challa6 min read
On this page
- Intro — what this post promises
- Lab topology
- Arm A — literal count (`shopper`)
- Arm B — find / search (do not trust early hits)
- Arm C — line prefix `ERROR`
- Arm D — case-insensitive + structured extracts
- How to read these numbers
- Pitfalls we hit (or avoided)
- Practical checklist
- Methodology footnote
- Ratio cheat sheet (time\_re / time\_str; >1 means re slower)
- Versions / environment pinned
- Verdict
Intro — what this post promises
Reach for re when you need patterns. Reach for str when you need speed on literals. This lab puts numbers on that slogan over a ~32 MiB synthetic log-ish corpus.
This is a hands-on lab with measured numbers:
- Literal count —
str.countvsre.findall/finditer. - find / search — including missing and late needles (full scans).
- Line prefix —
startswithscan vsre.compile(r'^ERROR', re.M). - Case-insensitive and “email-ish” / capture-group shapes.
Related links:
- json vs orjson vs msgpack localhost lab
- SHA-256 vs BLAKE2b vs xxHash localhost lab
- zstd vs gzip vs lz4 compression localhost lab
- stdout buffering line vs full localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (1 Oct 2026 IST): Shared Linux lab box. Python 3.13.5. Corpus ~32 MiB, ~395k lines, written to evidence. Affiliates: 0. Microbench on synthetic text — not production PCRE2, not a parser.
Verdict up front: str.count ~973 MB/s vs re.findall ~637 MB/s on a literal (1.53×). Prefix 4.6×). Simple capture extract nearly tied.ERROR lines: str scan ~1016 MB/s vs multiline re ~219 MB/s (
Lab topology
Script: lab-evidence/39-python-re-vs-str/results/run_lab.py.
Arm A — literal count (shopper)
| Method | p50 | MB/s | hits |
|---|---|---|---|
str.count | 32.9 ms | 973 | 409861 |
re.findall | 50.2 ms | 637 | 409861 |
re.finditer count | 59.0 ms | 542 | 409861 |
re.findall / str.count time ≈ 1.53×. Same hit count — fair fight.
Related links:
Arm B — find / search (do not trust early hits)
First occurrence of shopper sits near the start — both APIs return in microseconds; MB/s looks absurd. Full-scan arms:
| Method | p50 | MB/s | result |
|---|---|---|---|
str.find missing | 10.5 ms | 3039 | -1 |
re.search missing | 22.3 ms | 1438 | none |
str.find late needle | 7.0 ms | 4554 | EOF offset |
re.search late | 21.2 ms | 1509 | same |
Missing: re ~2.1× slower. Late: re ~3.0× slower. Use these, not the early-hit toy.
Arm C — line prefix ERROR
| Method | p50 | MB/s | lines |
|---|---|---|---|
splitlines + startswith | 65.1 ms | 492 | 7899 |
manual scan + startswith | 31.5 ms | 1016 | 7899 |
re.findall(r'^ERROR', re.M) | 146 ms | 219 | 7899 |
Avoid building a giant splitlines list when a scan suffices. Multiline ^ regex paid ~4.6× vs the scan.
Arm D — case-insensitive + structured extracts
| Method | p50 | MB/s |
|---|---|---|
text.lower().count('shopper') | 57.2 ms | 560 |
re.findall(..., re.I) | 337 ms | 94.9 |
| naive str email walker | 39.3 ms | 814 |
re.findall email pattern | 636 ms | 50.3 |
str scan GET /api/v1/items/<id> | 20.8 ms | 1541 |
re.findall with group | 22.8 ms | 1403 |
re.I ~5.9× slower than lower+count here (and lower allocates a full copy — still won). Email regex is more expressive; the ~16× gap is “pattern power costs CPU.” GET id extract is where a tight regex nearly matches a hand scanner.
Related links:
- json vs orjson vs msgpack localhost lab
- SHA-256 vs BLAKE2b vs xxHash localhost lab
- stdout buffering line vs full localhost lab
How to read these numbers
- Literals →
str. count/find/startswith are hard to beat. - Real patterns →
re. Pay the tax knowingly (emails, alternation, lookarounds). - Compile once. All re arms used
re.compile. - MB/s assumes a full corpus pass — early-exit finds need separate labeling.
Pitfalls we hit (or avoided)
- O(n²) corpus builder (
sum(len(line))each append) — hung the first run; fixed with a running counter. - Citing early
findMB/s as throughput — nonsense; added missing/late arms. - Comparing
splitlines+ startswith to regex without a scan baseline. - Calling naive email walker “equivalent” to RFC regex — same hit count here, different power.
- Forgetting
re.Ivs explicit.lower()allocation tradeoffs.
Practical checklist
- Hot path literal scans: prefer
strmethods. - Need a grammar: use
re, precompile, measure. - Prefer scanning over
splitlines()on large blobs when possible. - For casefold counting, bench
.lower()vsre.Ion your data. - Keep paired asserts so str/re refactors do not drift on counts.
Related links:
- process vs thread pool GIL localhost lab
- asyncio vs threads IO concurrency lab
- posix fadvise sequential vs random localhost lab
- fork COW RSS vs spawn localhost lab
Methodology footnote
Corpus mixes INFO filler, periodic ERROR lines, GET paths, and userN@example.com lines. Timed with time.perf_counter. MB/s = UTF-8 byte length / p50 seconds. Result counts asserted equal across paired arms (count, ERROR lines, GET ids).
Ratio cheat sheet (time_re / time_str; >1 means re slower)
| Task | Ratio |
|---|---|
| literal findall vs count | 1.53× |
| missing search vs find | 2.11× |
| late search vs find | 3.02× |
^ERROR vs scan startswith | ~4.6× |
re.I findall vs lower+count | ~5.9× |
| email findall vs naive walker | ~16× |
| GET id findall vs scan | ~1.10× |
When the ratio approaches 1.0, readability may win and you keep the regex. When it is 4×+ on a hot path, write the literal scanner (with tests) or narrow the regex.
For multi-gigabyte logs, consider streaming line-by-line (still str.startswith / compiled re on each line) so you do not hold a 32 MiB+ contiguous unicode object. Throughput will drop; peak RSS will thank you — see the fork COW lab for why large private buffers hurt.
Versions / environment pinned
- Python 3.13.5 stdlib
re(no third-party regex engine) - Corpus ~32 MiB / ~395k lines at
lab-evidence/39-python-re-vs-str/results/corpus.txt - Patterns precompiled; benches use p50 of 7 timed calls after warmup
Bottom line for Python services: default to str for fixed tokens and reserved words in logs; promote to re when the token shape is genuinely a language. The GET-id near-parity case shows a careful regex need not be slow — it also shows you should still verify against a scanner before declaring victory.
Verdict
On a ~32 MiB corpus, literal str.count (~973 MB/s) beat re.findall (~637 MB/s) by ~1.53×, and an ERROR-prefix scan beat multiline regex by ~4.6×. A simple GET-id regex nearly matched a hand scanner. Use str for literals; keep re for patterns — and measure the boundary.
Evidence path on the lab box: lab-evidence/39-python-re-vs-str/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5; ~32 MiB synthetic corpus (~395k lines). str.count('shopper') 973 MB/s vs re.findall 637 (~1.53x slower). ERROR line prefix: str scan 1016 MB/s vs re^ERROR 219 (~4.6x). Case-insensitive: str.lower().count 560 vs re.I 95 (~5.9x). Missing needle full scan: str.find 3039 MB/s vs re.search 1438 (~2.1x). Emails: naive str 814 vs re 50 (~16x; pattern power differs). GET id extract: str 1541 vs re 1403 (close). Affiliates: 0. Evidence: lab-evidence/39-python-re-vs-str/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026