Plate 22
str.translate vs replace: Scrub Lab
Hands-on str.translate vs replace vs re.sub lab: real ops/s for character delete and map scrubbing on mid-size strings, measured on Linux localhost (lab).
Aditya Challa4 min read
Intro — what this post promises
Scrubbing characters: is str.translate faster than a chained str.replace or re.sub? This lab times delete, map, and combined tables on mid-size strings on Linux localhost (numpy-free, stdlib only).
Related links:
- string concat vs join localhost lab
- bytes vs bytearray localhost lab
- setdefault vs defaultdict localhost lab
- zip vs index pairing localhost lab
- functools partial vs lambda localhost lab
- itertools chain vs flatten localhost lab
- statistics vs manual mean localhost lab
- perf_counter vs time localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Delete set ≈ punctuation + tab/newline; map = vowels + digits → markers. Affiliates: 0. Short strings can favor few replace calls; longer scrub favors translate.
Verdict up front (len=2 000 delete): translate ~317M ch/s vs replace-chain (~8.1×) vs re.sub (~8.9×). Combined delete+map: translate ~9.6× chained replace. Prefer str.maketrans + translate for multi-char scrub.
Arms
| Arm | Pattern |
|---|---|
| translate delete | maketrans('', '', delete_chars) |
| replace chain delete | s.replace(ch,'') per punct char |
re.sub delete | compiled character class |
| translate map | maketrans(from, to) |
| replace chain map | one replace per mapped char |
re.sub map | callback per match |
| translate combined | map + delete in one table |
Lab topology
Script: lab-evidence/73-str-translate-vs-replace/results/run_lab.py.
Lead table — delete (p50 chars/s)
| Arm | len=200 | len=2 000 | len=20 000 |
|---|---|---|---|
| translate | 63.9M | 317M | 245M |
| replace chain | 36.2M | 39.0M | 34.9M |
| re.sub | 34.4M | 35.7M | 31.6M |
translate÷replace delete k: ~8.1×; k: ~7.0×.
Map & combined (len=2 000)
| Arm | str/s | chars/s |
|---|---|---|
| replace chain map | 207.9k | 416M |
| translate map | 174.0k | 348M |
| re.sub map | 11.5k | 23.0M |
| translate combined | 172.2k | 344M |
| replace delete+map | 18.0k | 35.9M |
Honesty: with only 20 map chars, chained replace can beat translate map on short inputs (len=200: translate÷replace map ~0.32×). At len=20 000 translate map wins (~1.56×). Many deletes flip the story hard toward translate.
Reading it
- Many deletions / one pass —
translatewith a delete set is the winner (~8× replace at 2 k). - Few substitutions on tiny strings — a handful of
replacecalls can win; still prefer translate for maintainability as the set grows. re.subfor char classes — usually last for this job; keep regex for real patterns.- Build
maketransonce — amortize table construction outside the loop.
When a few replace calls are enough
If you only fold three characters, chained replace can beat a general translate map on short inputs — fewer passes than a wide punctuation scrub, and less table machinery. The crossover is about how many distinct edits you apply per string, not a blanket “replace is slow.” Once the delete/map set looks like “all punctuation” or “normalize a keyboard row,” switch to maketrans and stop growing the chain.
Pitfalls
- Chaining 50×
replacefor punctuation — quadratic-ish passes over the string. - Rebuilding translate tables every call — hoist
maketrans. - Using regex for single-char deletes — translate is purpose-built.
- Assuming replace is always slower — measure when the replace count is tiny.
When to pick what
| Need | Prefer |
|---|---|
| Delete/map many chars | str.translate |
| 1–3 fixed substitutions | str.replace |
| Patterns / unicode classes | re |
| Delete + map together | one maketrans(..., delete) |
Reproduce
Evidence: /workspace/lab-evidence/73-str-translate-vs-replace/results/.
Closing
translate for scrub tables; replace for tiny fixed edits; regex last. On this box len=2 k delete translate beat replace-chain by ~8.1× and re.sub by ~8.9×; combined scrub stayed ~9.6× ahead of chained replace. Build one table when the character set is non-trivial.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. N=2000 delete: translate 317M ch/s vs replace-chain ~8.1x vs re.sub ~8.9x; combined translate vs replace-both ~9.6x. Affiliates: 0. Evidence: lab-evidence/73-str-translate-vs-replace/.
Related links
Plate 71
html.escape vs Manual Replace: Localhost Lab
A hands-on localhost lab comparing html.escape with chained str.replace for safe HTML escaping.
Observability & SRE · 30 Sept 2026
Plate 18
fnmatch vs re Name Filter: Localhost Lab
Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.
Observability & SRE · 30 Sept 2026
Plate 84
tarfile vs zipfile Create+Extract: Localhost Lab
Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.
Observability & SRE · 30 Sept 2026