Plate 16
struct.pack vs to_bytes vs memoryview Lab
Benchmarking fixed-record packing with struct.pack, Struct.pack_into, int.to_bytes, and memoryview on Linux localhost.
Aditya Challa3 min read
Intro — what this post promises
Packing a fixed record — u32, u16, u16, u64 (16 bytes, little-endian) — with struct.pack, Struct.pack_into, int.to_bytes, or a memoryview into a bytearray. This lab reports records/s and MB/s on Linux localhost.
Related links:
- base64 vs hex encode localhost lab
- json vs orjson vs msgpack localhost lab
- string concat vs join localhost lab
- copy vs deepcopy localhost lab
- logging vs print localhost lab
- SHA-256 vs BLAKE2 vs xxHash localhost lab
- uuid4 vs secrets token localhost lab
- atomic rename vs overwrite localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. 200,000 records per batch. Affiliates: 0. No C extension beyond stdlib struct.
Verdict up front: Struct.pack_into ~125.7 MB/s (8.24M rec/s, 121 ns). struct.pack + b''.join is ~74.7 MB/s. to_bytes + memoryview is ~35.2 MB/s (~0.47× pack). Hand-rolled shifts are the slowest (~14.5 MB/s).
Layout
Arms: struct.pack then join; compiled Struct.pack then join; pack_into a pre-sized bytearray; four to_bytes pieces joined; to_bytes written through a memoryview; manual byte shifts into a bytearray.
Script: lab-evidence/53-struct-pack-vs-to-bytes/results/run_lab.py.
Lead table — p50
| Arm | rec/s | MB/s | ns/rec | vs struct.pack |
|---|---|---|---|---|
| pack_into | 8,237,581 | 125.7 | 121.4 | 1.68× |
| Struct.pack + join | 5,302,553 | 80.9 | 188.6 | 1.08× |
| struct.pack + join | 4,893,340 | 74.7 | 204.4 | 1.00× |
| to_bytes + memoryview | 2,303,908 | 35.2 | 434.0 | 0.47× |
| to_bytes + join | 1,993,448 | 30.4 | 501.6 | 0.41× |
| manual shifts | 953,216 | 14.5 | 1049.1 | 0.19× |
Reading it
pack_intowins because it does not allocate a 16-bytebytesper record and then join them. One buffer, one write cursor.Structvsstruct.packis a small win (80.9 vs 74.7 MB/s) — the format is compiled once. Still slower thanpack_intobecause of the per-record object + join.int.to_bytesbuilds four temporarybytesper record. memoryview avoids the final join but not those temps: 35.2 vs 30.4 MB/s.- Manual shifts look “closer to the metal” and lose (1049 ns/rec). Python bytecode per byte is the tax;
structis C.
Pitfalls
- Timing
packwithout the join — you still have to assemble the buffer in real code. - Native endian (
@/=) — this lab is explicit<. - Assuming memoryview makes
to_bytesfree — the allocations are insideto_bytes. - One-off headers — clarity of
to_bytesis fine; this gap shows up at hundreds of thousands of records.
When to pick what
| Need | Prefer |
|---|---|
| Many fixed records into one buffer | Struct.pack_into |
| A handful of fields, readable code | struct.pack or to_bytes |
| Already have a writable buffer | pack_into / memoryview |
| “I’ll just shift bits” | measure first — here it was slowest |
Reproduce
Evidence: /workspace/lab-evidence/53-struct-pack-vs-to-bytes/results/.
Closing
Write into a buffer. On this box pack_into ~126 MB/s beats struct.pack+join ~75 and to_bytes ~30–35. Manual shifts landed at ~15 MB/s. Use struct for layouts; use pack_into when the record rate is the point.
Lab evidence
What I found running this
Ran python3 lab-evidence/53-struct-pack-vs-to-bytes/results/run_lab.py on 1 Oct 2026 IST with Python 3.13.5 on Linux localhost. Measured 200,000 fixed 16-byte records across struct.pack plus join, Struct.pack plus join, Struct.pack_into, to_bytes plus join, to_bytes through memoryview, and manual shifts. p50 results: pack_into 8.24M records/s (125.7 MB/s), Struct.pack plus join 80.9 MB/s, struct.pack plus join 74.7 MB/s, to_bytes plus memoryview 35.2 MB/s, to_bytes plus join 30.4 MB/s, and manual shifts 14.5 MB/s. The surprising result was that memoryview avoids the final join but cannot remove to_bytes allocations.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026