ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 41

  1. Blog
  2. /Observability & SRE

shutil.copyfile vs Manual Copy Lab

Hands-on shutil.copyfile vs manual file copy lab: real MB/s versus open/read/write and Path.read_bytes on local disk, measured on Linux localhost (lab).

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — throughput (p50 MB/s)
  5. Small-file copies/s (4 KiB)
  6. Reading it
  7. Why not pathlib join again?
  8. copyfile vs copy2 in one line
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing

Intro — what this post promises

Copying file bytes on disk: is shutil.copyfile faster than a manual open + read/write loop (or Path.read_bytes / write_bytes)? This lab measures MB/s on Linux localhost temp files — not another pathlib-vs-os.path join lab (that is already published).

Related links:

  • pathlib vs os.path localhost lab
  • sendfile vs userspace copy localhost lab
  • bytes vs bytearray localhost lab
  • str translate vs replace localhost lab
  • setdefault vs defaultdict localhost lab
  • copy vs deepcopy localhost lab
  • json dumps compact vs indent localhost lab
  • perf_counter vs time localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Same filesystem tempdir. Affiliates: 0. copyfile copies data only; copy2 also copies metadata.

Verdict up front: 1 MiB — copyfile ~1823 MB/s vs manual 64 KiB chunks ~1542 (~1.18×) vs Path read/write ~1557. 8 MiB — copyfile ~1346 vs read-all ~879 (~1.53×); chunked manual stayed close. Prefer shutil.copyfile (or copy2 when you need mtime/mode); avoid slurping large files into RAM.


Arms

ArmPattern
copyfileshutil.copyfile(src, dst)
copy2data + metadata
manual chunk 64 KiBread/write loop
manual readallread() then write()
Path read/writedst.write_bytes(src.read_bytes())

Sizes: 4 KiB / 64 KiB / 1 MiB / 8 MiB with repeated copies.


Lab topology

tempfile dir on local Linux disk
metric: p50 MB/s and copies/s

Script: lab-evidence/74-shutil-copyfile-vs-manual/results/run_lab.py.


Lead table — throughput (p50 MB/s)

Arm4 KiB64 KiB1 MiB8 MiB
copyfile6163418231346
copy25455117381931
manual chunk7273015421420
manual readall737331549879
Path r/w6972215571394

Small-file copies/s (4 KiB)

Armcopies/s
manual readall18,724
manual chunk18,550
Path r/w17,637
copyfile15,720
copy213,845

Tiny files are dominated by open/unlink overhead — manual can look “faster” in a tight loop.


Reading it

  • MiB-class payloads — copyfile edged chunked manual by ~1.18× here (C helper, fewer Python trips).
  • Don’t read() entire large files — at 8 MiB readall trailed copyfile by ~1.53×.
  • copy2 costs metadata on small files; on larger buffered copies it can still win on this box (cache/FS noise — treat as “same league”).
  • Path convenience ≈ manual readall/chunk for mid sizes — fine for scripts, not a speedup.

Why not pathlib join again?

Lab 47 already covered Path vs os.path for join/exists/stat. This lab answers a different question: how you move bytes once you have two paths.


copyfile vs copy2 in one line

Need only bytes? copyfile. Need the destination to look like the source in ls -l timestamps/mode? copy2. Do not invent a third wrapper that stats and chmods by hand unless you are filtering metadata on purpose.


Pitfalls

  1. Hand-rolling copies “for control” without chunking large files.
  2. Using copy2 when you only need bytes — extra syscalls on small files.
  3. Benchmarking only 4 KiB — open overhead hides copy algorithm.
  4. Cross-filesystem / remote mounts — numbers here are local tempdir.

When to pick what

NeedPrefer
Data-only copyshutil.copyfile
Preserve mtime/modeshutil.copy2
Custom transform while copyingmanual chunked loop
Tiny script conveniencePath read/write

Reproduce

python3 lab-evidence/74-shutil-copyfile-vs-manual/results/run_lab.py

Evidence: /workspace/lab-evidence/74-shutil-copyfile-vs-manual/results/.


Closing

Use shutil for whole-file copies; chunk if you must DIY. On this box 1 MiB copyfile hit ~1823 MB/s (~1.18× chunked manual), and at 8 MiB it beat full-file readall by ~1.53×. Reach for pathlib when paths are the story — for bytes, copyfile is the default.

shutil.copyfileshutil.copy2file copyread writepythonlocalhost labsrecopyfile

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. Same filesystem tempdir. Measured 4 KiB, 64 KiB, 1 MiB, and 8 MiB repeated copies. At 1 MiB copyfile reached ~1823 MB/s vs manual chunk ~1542; at 8 MiB copyfile ~1346 vs readall ~879. Affiliates: 0. Evidence: lab-evidence/74-shutil-copyfile-vs-manual/results/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 18

    fnmatch vs re Name Filter: Localhost Lab

    Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.

    Observability & SRE · 30 Sept 2026

  • Plate 84

    tarfile vs zipfile Create+Extract: Localhost Lab

    Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

  • Plate 85

    scandir vs listdir vs iterdir: Localhost Lab

    Hands-on os.scandir vs listdir vs Path.iterdir lab: real entries/s for names and is_file on a synthetic tree, measured on Linux localhost (lab) for SREs.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — throughput (p50 MB/s)
  5. Small-file copies/s (4 KiB)
  6. Reading it
  7. Why not pathlib join again?
  8. copyfile vs copy2 in one line
  9. Pitfalls
  10. When to pick what
  11. Reproduce
  12. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove