ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 17

  1. Blog
  2. /Observability & SRE

difflib vs set Ops Similarity: Localhost Lab

Hands-on difflib.SequenceMatcher vs set Jaccard token similarity: real ops/s on token lists, measured on Linux localhost in this hands-on lab for SREs.

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 800 tokens (p50)
  5. Scale sketch (ops/s)
  6. Which tool when
  7. Order vs bag
  8. Reading it
  9. Multiplicity caveat
  10. Opcode consumers
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway

Intro — what this post promises

Compare two moderately different token lists: difflib.SequenceMatcher (.ratio() / .get_opcodes()) vs set intersection/union Jaccard. This lab reports ops/s on Linux localhost.

Frame: difflib for edit-aware similarity and opcodes; sets for bag-of-tokens overlap. Affiliates: 0.

Related links:

  • zlib vs gzip compress localhost lab
  • queue vs deque handoff localhost lab
  • stringio vs list join localhost lab
  • groupby vs manual localhost lab
  • methodcaller vs getattr localhost lab
  • hmac compare digest localhost lab
  • argparse vs sys argv localhost lab
  • shelve vs pickle dict localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. No Docker. Tokens drawn from a 38-word vocab; mutate ≈15% replace / 5% delete / 5% insert.

Verdict up front (800 tokens): set Jaccard ~104460.5 ops/s; SequenceMatcher.ratio (tokens) ~4517.3; opcodes ~4410.9; ratio on joined strings ~33.8. Use sets for overlap; difflib when you need edit ops.


Arms

ArmWhat it answers
SequenceMatcher(...).ratio() on token listsedit-aware similarity
.get_opcodes()insert/delete/replace spans
.ratio() on joined stringschar-level (costly)
set Jaccardtoken overlap
inter/union counts onlyraw set ops

Lab topology

n tokens: 200 · 800 · 2000 · 7 rounds · p50
metric: ops/s = 1 / p50_s

Script: lab-evidence/103-difflib-vs-set-ops/results/run_lab.py.


Lead table — 800 tokens (p50)

Armops/snote
set Jaccard104460.5jaccard=1.000
set inter∪union107296.1counts only
SequenceMatcher.ratio (tokens)4517.3ratio=0.003
get_opcodes (tokens)4410.92 opcodes
ratio on joined str33.8char sequence

Sets win by ~23× vs token-level SequenceMatcher here — different questions, different asymptotics.


Scale sketch (ops/s)

nset JaccardSM ratio (tokens)SM ratio (joined str)
200249563.26863.7518.5
800104460.54517.333.8
200037267.51788.92.7

Joined-string ratio() collapses as character length grows — tokenize first if you only care about words.


Which tool when

  • Jaccard / set ops: duplicate detection, tag overlap, “same bag of words?” — fast and simple.
  • SequenceMatcher: unified diffs, change highlighting, edit distance–ish ratios when order matters.
  • Opcodes: drive a UI that shows insert/delete/replace spans.

Do not use Jaccard as a drop-in for edit similarity — reordered identical tokens score high on sets and lower on difflib.


Order vs bag

If two lines swap token order, Jaccard may stay high while SequenceMatcher.ratio drops. Pick the metric that matches user-visible “sameness.” For changelog UIs, opcodes beat any scalar overlap score.


Reading it

  • Tokenize before SequenceMatcher when inputs are prose-sized.
  • Prefer sets when multiplicity/order do not matter.
  • Cache SequenceMatcher results if you call both ratio and opcodes (autojunk / internals).
  • Report the metric that matches the product question.

Multiplicity caveat

set Jaccard collapses duplicate tokens. If “three copies of error” should weigh more than one, use collections.Counter multiset overlap instead — slower than pure sets, still usually far cheaper than SequenceMatcher on long sequences.


Opcode consumers

get_opcodes() returns 5-tuples (tag, i1, i2, j1, j2). Driving a diff UI from opcodes is the usual win; calling ratio() alone when you already need opcodes wastes a second pass unless you reuse one matcher instance.


Pitfalls

  • Running SequenceMatcher on multi‑KB raw strings in a hot loop.
  • Treating Jaccard == SequenceMatcher.ratio numerically (they diverge).
  • Ignoring duplicates — sets collapse multiplicity; bags need Counter.
  • Comparing autojunk on/off without documenting it.

Reproduce

python3 lab-evidence/103-difflib-vs-set-ops/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box. Synthetic vocab + seeded mutations. Not difflib.unified_diff file I/O. Not embedding cosine similarity.


Takeaway

At 800 tokens, set Jaccard ~104460.5 ops/s crushed SequenceMatcher.ratio ~4517.3 ops/s. Use sets for overlap, difflib for edit ops — and avoid char-level ratio() on long joined strings (~33.8 ops/s here).

difflib.sequencematcherjaccardset intersectiontoken similarityget_opcodeslocalhost labsreops/s

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. 800 tokens: set Jaccard 104460.5 ops/s; SM ratio tokens 4517.3; opcodes 4410.9; joined-str ratio 33.8. Affiliates: 0. Evidence: lab-evidence/103-difflib-vs-set-ops/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 71

    html.escape vs Manual Replace: Localhost Lab

    A hands-on localhost lab comparing html.escape with chained str.replace for safe HTML escaping.

    Observability & SRE · 30 Sept 2026

  • Plate 63

    Decimal vs float Sum: Localhost Lab

    Hands-on Decimal vs float cumulative sum lab: real ops/s plus a simple accuracy note (not financial advice), measured on Linux localhost (lab) for SREs.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — 800 tokens (p50)
  5. Scale sketch (ops/s)
  6. Which tool when
  7. Order vs bag
  8. Reading it
  9. Multiplicity caveat
  10. Opcode consumers
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove