Plate 17
difflib vs set Ops Similarity: Localhost Lab
Hands-on difflib.SequenceMatcher vs set Jaccard token similarity: real ops/s on token lists, measured on Linux localhost in this hands-on lab for SREs.
Aditya Challa4 min read
Intro — what this post promises
Compare two moderately different token lists: difflib.SequenceMatcher (.ratio() / .get_opcodes()) vs set intersection/union Jaccard. This lab reports ops/s on Linux localhost.
Frame: difflib for edit-aware similarity and opcodes; sets for bag-of-tokens overlap. Affiliates: 0.
Related links:
- zlib vs gzip compress localhost lab
- queue vs deque handoff localhost lab
- stringio vs list join localhost lab
- groupby vs manual localhost lab
- methodcaller vs getattr localhost lab
- hmac compare digest localhost lab
- argparse vs sys argv localhost lab
- shelve vs pickle dict localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. No Docker. Tokens drawn from a 38-word vocab; mutate ≈15% replace / 5% delete / 5% insert.
Verdict up front (800 tokens): set Jaccard ~104460.5 ops/s; SequenceMatcher.ratio (tokens) ~4517.3; opcodes ~4410.9; ratio on joined strings ~33.8. Use sets for overlap; difflib when you need edit ops.
Arms
| Arm | What it answers |
|---|---|
SequenceMatcher(...).ratio() on token lists | edit-aware similarity |
.get_opcodes() | insert/delete/replace spans |
.ratio() on joined strings | char-level (costly) |
set Jaccard | token overlap |
| inter/union counts only | raw set ops |
Lab topology
Script: lab-evidence/103-difflib-vs-set-ops/results/run_lab.py.
Lead table — 800 tokens (p50)
| Arm | ops/s | note |
|---|---|---|
| set Jaccard | 104460.5 | jaccard=1.000 |
| set inter∪union | 107296.1 | counts only |
| SequenceMatcher.ratio (tokens) | 4517.3 | ratio=0.003 |
| get_opcodes (tokens) | 4410.9 | 2 opcodes |
| ratio on joined str | 33.8 | char sequence |
Sets win by ~23× vs token-level SequenceMatcher here — different questions, different asymptotics.
Scale sketch (ops/s)
| n | set Jaccard | SM ratio (tokens) | SM ratio (joined str) |
|---|---|---|---|
| 200 | 249563.2 | 6863.7 | 518.5 |
| 800 | 104460.5 | 4517.3 | 33.8 |
| 2000 | 37267.5 | 1788.9 | 2.7 |
Joined-string ratio() collapses as character length grows — tokenize first if you only care about words.
Which tool when
- Jaccard / set ops: duplicate detection, tag overlap, “same bag of words?” — fast and simple.
- SequenceMatcher: unified diffs, change highlighting, edit distance–ish ratios when order matters.
- Opcodes: drive a UI that shows insert/delete/replace spans.
Do not use Jaccard as a drop-in for edit similarity — reordered identical tokens score high on sets and lower on difflib.
Order vs bag
If two lines swap token order, Jaccard may stay high while SequenceMatcher.ratio drops. Pick the metric that matches user-visible “sameness.” For changelog UIs, opcodes beat any scalar overlap score.
Reading it
- Tokenize before SequenceMatcher when inputs are prose-sized.
- Prefer sets when multiplicity/order do not matter.
- Cache
SequenceMatcherresults if you call bothratioandopcodes(autojunk / internals). - Report the metric that matches the product question.
Multiplicity caveat
set Jaccard collapses duplicate tokens. If “three copies of error” should weigh more than one, use collections.Counter multiset overlap instead — slower than pure sets, still usually far cheaper than SequenceMatcher on long sequences.
Opcode consumers
get_opcodes() returns 5-tuples (tag, i1, i2, j1, j2). Driving a diff UI from opcodes is the usual win; calling ratio() alone when you already need opcodes wastes a second pass unless you reuse one matcher instance.
Pitfalls
- Running SequenceMatcher on multi‑KB raw strings in a hot loop.
- Treating Jaccard == SequenceMatcher.ratio numerically (they diverge).
- Ignoring duplicates — sets collapse multiplicity; bags need Counter.
- Comparing autojunk on/off without documenting it.
Reproduce
Evidence: summary.json, summary.txt.
Limits
One Linux box. Synthetic vocab + seeded mutations. Not difflib.unified_diff file I/O. Not embedding cosine similarity.
Takeaway
At 800 tokens, set Jaccard ~104460.5 ops/s crushed SequenceMatcher.ratio ~4517.3 ops/s. Use sets for overlap, difflib for edit ops — and avoid char-level ratio() on long joined strings (~33.8 ops/s here).
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. 800 tokens: set Jaccard 104460.5 ops/s; SM ratio tokens 4517.3; opcodes 4410.9; joined-str ratio 33.8. Affiliates: 0. Evidence: lab-evidence/103-difflib-vs-set-ops/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 71
html.escape vs Manual Replace: Localhost Lab
A hands-on localhost lab comparing html.escape with chained str.replace for safe HTML escaping.
Observability & SRE · 30 Sept 2026
Plate 63
Decimal vs float Sum: Localhost Lab
Hands-on Decimal vs float cumulative sum lab: real ops/s plus a simple accuracy note (not financial advice), measured on Linux localhost (lab) for SREs.
Observability & SRE · 30 Sept 2026