Plate 32
array.array vs list[int] vs bytes Lab
A hands-on Linux localhost lab comparing memory density and numeric throughput for list[int], array.array('i'), bytearray, and memoryview.
Aditya Challa4 min read
Intro — what this post promises
Storing millions of integers: list[int], array.array('i'), or a packed bytearray / memoryview? Folklore says arrays are “more compact and faster.” This lab measures bytes per element and sum / iterate / index ops/s on Linux localhost.
Related links:
- deque vs list queue localhost lab
- itertools vs python loops localhost lab
- dataclass vs slots vs dict localhost lab
- struct pack vs to_bytes localhost lab
- copy vs deepcopy localhost lab
- lru_cache hit vs miss localhost lab
- set vs list membership localhost lab
- json vs orjson vs msgpack localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. N=2,000,000. Affiliates: 0. Density uses itemsize / estimated sys.getsizeof — not a full heap profiler. sum(list) is a C-accelerated special case.
Verdict up front: density wins for array('i') — 4 B/elem vs list ~36 B/elem (~9×). Throughput is the twist: sum(list) ~196M/s beats sum(array) ~102M because each C int becomes a Python int. Manual bytearray unpack is ~5.8M/s.
Structures
| Kind | Layout |
|---|---|
list(range(N)) | pointers to PyLong objects |
array.array('i', …) | contiguous signed 32-bit |
bytearray of tobytes() | raw little-endian i32 |
memoryview(ba).cast('i') | typed view over bytes |
Lab topology
Script: lab-evidence/56-array-vs-list-ints/results/run_lab.py.
Memory density
| Structure | bytes / elem (approx) |
|---|---|
array('i') | 4 |
bytearray (i32 pack) | 4 |
list[int] | ~36 |
sys.getsizeof on the list object alone is tiny; the ~36 B figure includes a sample of int object sizes × N. That is the RAM story for large numeric vectors.
Lead table — throughput (p50)
| Arm | ops/s | ns/op |
|---|---|---|
| sum(list) | 195,746,143 | 5.1 |
| sum(array 'i') | 102,095,834 | 9.8 |
| sum(memoryview 'i') | 88,346,734 | 11.3 |
| sum(bytearray unpack) | 5,808,188 | 172.2 |
| for+= list | 52,402,418 | 19.1 |
| for+= array | 37,159,104 | 26.9 |
| for+= memoryview | 35,689,755 | 28.0 |
| index stride7 list | 24,480,313 | 40.8 |
| index stride7 array | 24,558,484 | 40.7 |
| index stride7 mv | 23,070,925 | 43.3 |
Reading it
- RAM:
array('i')/ packed bytes are ~9× denser than a list of small ints. That alone can decide caching and GC pressure. sum(list)is not a fair “list is faster” lesson — CPython’ssumon lists of ints is specialized. Fairer: thefor x in …: s += xarms, where list is still ahead (~1.41×) because array iteration boxes.- Index stride is nearly tied — boxing happens either way when you pull a Python int out.
- Hand-unpacking bytearray is for I/O buffers, not inner-loop math.
For numeric crunching that stays in C (NumPy), you leave Python’s boxing tax entirely. This lab is stdlib only.
Pitfalls
- Choosing
arrayfor speed ofsum()— memory was the win;sum(list)can win. array('l')/ platform sizes —'i'is 4 bytes here; checkitemsize.- Assuming memoryview is free — iteration still produces Python ints.
- RSS peak noise — prefer buffer nbytes for density claims.
When to pick what
| Need | Prefer |
|---|---|
| Millions of ints, RAM-bound | array.array / packed buffer |
| App logic, small N | list |
| Wire / file bytes | bytearray / memoryview |
| Heavy numeric compute | NumPy / similar (out of scope) |
Reproduce
Evidence: /workspace/lab-evidence/56-array-vs-list-ints/results/.
Closing
Compact ≠ always faster in pure Python. On this box array('i') is 4 B/elem vs list ~36 B, but sum(list) hits ~196M/s vs array ~102M. Use arrays for memory; don’t expect them to beat list iteration without staying in C.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Ran Python 3.13.5 with N=2,000,000. Measured array('i') at 4 B/elem, list at about 36 B/elem, and bytearray i32 at 4 B/elem. sum(): list 196M/s, array 102M/s, memoryview 88M/s, bytearray unpack 5.8M/s; iter+= list 52M/s vs array 37M/s. Affiliates: 0. Evidence: lab-evidence/56-array-vs-list-ints/.
Related links
Plate 87
mmap vs read Byte-Sum Scan: Localhost Lab
Hands-on mmap vs read vs chunked sequential byte-sum scan lab: real MB/s on a generated fixture (not page-touch), measured on Linux localhost for SREs.
30 Sept 2026
Plate 72
bytes vs bytearray: Mutate/Copy Lab
Hands-on bytes vs bytearray lab: real ops/s for append/extend/slice/copy and when copy dominates over mutate, benchmarked on Linux localhost for SREs.
30 Sept 2026
Plate 16
struct.pack vs to_bytes vs memoryview Lab
Benchmarking fixed-record packing with struct.pack, Struct.pack_into, int.to_bytes, and memoryview on Linux localhost.
Observability & SRE · 30 Sept 2026