ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 68

  1. Blog
  2. /Observability & SRE

gc.collect Cost Empty vs Cycles: Localhost Lab

Hands-on gc.collect cost empty vs cyclic garbage lab: real collect latency plus reclaim counts, measured on Linux localhost today in this lab for SREs.

Aditya Challa·1 October 2026·4 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table
  5. Gen0 is not enough
  6. Reading it for SRE work
  7. Allocation context
  8. Operational takeaway for latency budgets
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway

Intro — what this post promises

Cost of gc.collect() when the heap is empty vs when it holds cyclic garbage. This lab reports collect latency, collects/s, and objects reclaimed on Linux localhost.

Related links:

  • sqlite3 vs shelve localhost lab
  • fractions vs float localhost lab
  • math fsum vs sum localhost lab
  • itertools batched vs chunk localhost lab
  • exitstack vs nested with localhost lab
  • xml etree vs json localhost lab
  • logging formatter vs fstring localhost lab
  • copy copy vs dict copy localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. GC stayed enabled so cycles are tracked (disabled-GC setups under-count reclaim).

Verdict up front (50000 cycle pairs): empty full collect ~0.709 ms (~1411 collects/s); with cycles ~12.587 ms reclaiming ~100006 objects.


Arms

ArmPattern
full collect emptygc.collect()
gen0 emptygc.collect(0)
full with cyclesbuild 2-cycles, then full collect
gen0 after cyclesoften partial
full after gen0reclaim remainder
list alloc no cyclesallocation context

Seven rounds, p50. Each with-cycles arm rebuilds garbage after a clean collect.


Lab topology

n_cycle_pairs=50000 · 7 rounds · p50
metrics: p50_s, collects/s, p50_collected

Script: lab-evidence/129-gc-collect-cost/results/run_lab.py.


Lead table

Armp50 mscollects/scollected
full empty0.70914110
gen0 empty0.035460800
full + cycles12.58779100006
gen0 after cycles0.15564711964
full after gen012.3098198042

Gen0 is not enough

After building cycles, gen0 reclaimed only ~1964 objects; the following full collect took ~12.309 ms and reclaimed ~98042. Calling gc.collect(0) in a request path can look “cheap” while leaving cyclic junk behind.


Reading it for SRE work

  • Empty full collect is cheap here (~0.709 ms) — do not fear an occasional manual collect in a maintenance task.
  • Cyclic graphs (caches with parent/child links) make full collect ~17.8× slower than empty on this run.
  • Prefer breaking cycles (weakref, explicit clear) over hoping gen0 saves you.
  • Pair with RSS/tracemalloc labs if you later measure allocator pressure.

Allocation context

Building a plain 50000-int list (no cycles) ran at ~80686416 objs/s. That arm is not a collect — it shows allocation itself is far cheaper than a full collect over ~100006 cyclic objects.



Operational takeaway for latency budgets

A maintenance gc.collect() on a quiet process cost ~0.709 ms here. The same call after 50000 orphaned cycles jumped to ~12.587 ms. Put cycle-prone structures behind weakrefs or explicit teardown in long-lived workers; reserve full collects for controlled windows, not per-request cleanup.


Pitfalls

  • Disabling GC while creating objects, then wondering why collect() reclaims 0.
  • Using only gc.collect(0) as a “GC health check”.
  • Comparing collect latency across machines without stating object graph shape.
  • Forcing full collect in the hot path without a latency budget.

Reproduce

python3 lab-evidence/129-gc-collect-cost/results/run_lab.py

Evidence: summary.json, summary.txt.


Limits

One Linux box, CPython cyclic GC only. Not pymalloc arena tuning, not Go/Java collectors.


Manual collects belong in runbooks with a stated latency budget — not as a silent fix for “memory looked high” without graph evidence.


Takeaway

Empty gc.collect() ~0.709 ms; with 50000 cycle pairs ~12.587 ms reclaiming ~100006 objects. Prefer breaking cycles; do not trust gen0 alone for cyclic heaps.

pythongarbage collectiongc.collectmemory managementperformancelatencycyclic referencesbenchmarking

Lab evidence

What I found running this

Ran the gc.collect benchmark on Linux localhost with Python 3.13.5, seven rounds at p50, keeping GC enabled. Empty full collect measured about 0.709 ms; 50,000 cycle pairs took about 12.587 ms and reclaimed about 100006 objects. gen0 reclaimed only about 1964 first, which confirmed that a full collect is needed for cyclic garbage.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 16

    ExitStack vs Nested with Resources: Localhost Lab

    Hands-on contextlib.ExitStack vs nested with and manual close: real cycles/s for N resources, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 40

    itertools.batched vs Manual Chunking: Localhost Lab

    Hands-on itertools.batched vs manual list-slice chunking: real items/s batching sequences, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 28

    sqlite3 vs shelve Local KV: Localhost Lab

    Hands-on sqlite3 vs shelve local KV store lab: real insert/get ops/s plus file sizes, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table
  5. Gen0 is not enough
  6. Reading it for SRE work
  7. Allocation context
  8. Operational takeaway for latency budgets
  9. Pitfalls
  10. Reproduce
  11. Limits
  12. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove