ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 22

  1. Blog
  2. /Observability & SRE

str.translate vs replace: Scrub Lab

Hands-on str.translate vs replace vs re.sub lab: real ops/s for character delete and map scrubbing on mid-size strings, measured on Linux localhost (lab).

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — delete (p50 chars/s)
  5. Map & combined (len=2 000)
  6. Reading it
  7. When a few `replace` calls are enough
  8. Pitfalls
  9. When to pick what
  10. Reproduce
  11. Closing

Intro — what this post promises

Scrubbing characters: is str.translate faster than a chained str.replace or re.sub? This lab times delete, map, and combined tables on mid-size strings on Linux localhost (numpy-free, stdlib only).

Related links:

  • string concat vs join localhost lab
  • bytes vs bytearray localhost lab
  • setdefault vs defaultdict localhost lab
  • zip vs index pairing localhost lab
  • functools partial vs lambda localhost lab
  • itertools chain vs flatten localhost lab
  • statistics vs manual mean localhost lab
  • perf_counter vs time localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Delete set ≈ punctuation + tab/newline; map = vowels + digits → markers. Affiliates: 0. Short strings can favor few replace calls; longer scrub favors translate.

Verdict up front (len=2 000 delete): translate ~317M ch/s vs replace-chain (~8.1×) vs re.sub (~8.9×). Combined delete+map: translate ~9.6× chained replace. Prefer str.maketrans + translate for multi-char scrub.


Arms

ArmPattern
translate deletemaketrans('', '', delete_chars)
replace chain deletes.replace(ch,'') per punct char
re.sub deletecompiled character class
translate mapmaketrans(from, to)
replace chain mapone replace per mapped char
re.sub mapcallback per match
translate combinedmap + delete in one table

Lab topology

lengths: 200 / 2000 / 20000; many trials on short strings
ops = full-string transforms/s; also chars/s

Script: lab-evidence/73-str-translate-vs-replace/results/run_lab.py.


Lead table — delete (p50 chars/s)

Armlen=200len=2 000len=20 000
translate63.9M317M245M
replace chain36.2M39.0M34.9M
re.sub34.4M35.7M31.6M

translate÷replace delete  k: ~8.1×;  k: ~7.0×.


Map & combined (len=2 000)

Armstr/schars/s
replace chain map207.9k416M
translate map174.0k348M
re.sub map11.5k23.0M
translate combined172.2k344M
replace delete+map18.0k35.9M

Honesty: with only 20 map chars, chained replace can beat translate map on short inputs (len=200: translate÷replace map ~0.32×). At len=20 000 translate map wins (~1.56×). Many deletes flip the story hard toward translate.


Reading it

  • Many deletions / one pass — translate with a delete set is the winner (~8× replace at 2 k).
  • Few substitutions on tiny strings — a handful of replace calls can win; still prefer translate for maintainability as the set grows.
  • re.sub for char classes — usually last for this job; keep regex for real patterns.
  • Build maketrans once — amortize table construction outside the loop.

When a few replace calls are enough

If you only fold three characters, chained replace can beat a general translate map on short inputs — fewer passes than a wide punctuation scrub, and less table machinery. The crossover is about how many distinct edits you apply per string, not a blanket “replace is slow.” Once the delete/map set looks like “all punctuation” or “normalize a keyboard row,” switch to maketrans and stop growing the chain.


Pitfalls

  1. Chaining 50× replace for punctuation — quadratic-ish passes over the string.
  2. Rebuilding translate tables every call — hoist maketrans.
  3. Using regex for single-char deletes — translate is purpose-built.
  4. Assuming replace is always slower — measure when the replace count is tiny.

When to pick what

NeedPrefer
Delete/map many charsstr.translate
1–3 fixed substitutionsstr.replace
Patterns / unicode classesre
Delete + map togetherone maketrans(..., delete)

Reproduce

python3 lab-evidence/73-str-translate-vs-replace/results/run_lab.py

Evidence: /workspace/lab-evidence/73-str-translate-vs-replace/results/.


Closing

translate for scrub tables; replace for tiny fixed edits; regex last. On this box len=2 k delete translate beat replace-chain by ~8.1× and re.sub by ~8.9×; combined scrub stayed ~9.6× ahead of chained replace. Build one table when the character set is non-trivial.

str.translatemaketransstr.replacere.subpythonlocalhost labsrescrub

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. N=2000 delete: translate 317M ch/s vs replace-chain ~8.1x vs re.sub ~8.9x; combined translate vs replace-both ~9.6x. Affiliates: 0. Evidence: lab-evidence/73-str-translate-vs-replace/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 71

    html.escape vs Manual Replace: Localhost Lab

    A hands-on localhost lab comparing html.escape with chained str.replace for safe HTML escaping.

    Observability & SRE · 30 Sept 2026

  • Plate 18

    fnmatch vs re Name Filter: Localhost Lab

    Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.

    Observability & SRE · 30 Sept 2026

  • Plate 84

    tarfile vs zipfile Create+Extract: Localhost Lab

    Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — delete (p50 chars/s)
  5. Map & combined (len=2 000)
  6. Reading it
  7. When a few `replace` calls are enough
  8. Pitfalls
  9. When to pick what
  10. Reproduce
  11. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove