NumPy内存效率疑问:为何sys.getsizeof计算结果比Python列表高?
sys.getsizeof(p[3])*len(p) Shows Higher Memory for NumPy Array Than Python List? Great question! Let’s unpack this confusion step by step—this is a common gotcha when comparing NumPy arrays to Python lists.
First, let’s recap how each structure stores data:
- Python Lists: A list holds pointers to individual Python objects (even for integers). When you call
sys.getsizeof(l[3]), you’re getting the size of a single Pythonintobject (which includes overhead like reference counts and type info). Multiplying by the list length gives you the total size of all those individualintobjects (though it doesn’t include the list’s own internal overhead, but that’s not the core issue here). - NumPy Arrays: A NumPy array stores raw, homogeneous data in a single contiguous block of memory. It doesn’t store Python objects—instead, it stores primitive values (like 64-bit integers) directly. That’s why
p.itemsize * p.sizegives you the actual memory used by the array:itemsizeis the byte size of one primitive value, andsizeis the number of elements.
Now, why does sys.getsizeof(p[3])*len(p) give a higher number?
When you index into a NumPy array (like p[3]), NumPy doesn’t return the raw primitive value—it wraps it in a NumPy scalar object (e.g., numpy.int64). This object is a Python-compatible wrapper that lets you use the value like a regular Python int, but it comes with its own overhead (just like any Python object).
So sys.getsizeof(p[3]) is measuring the size of this wrapper object, not the raw 8-byte integer stored in the array. When you multiply that wrapper size by the number of elements, you’re essentially calculating what it would cost to store every element as a separate NumPy scalar object—which is way more than both the list’s int objects and the array’s raw data.
Example numbers to make it concrete:
Let’s say on your system:
sys.getsizeof(l[3])= 28 bytes (size of a Pythonintobject for small integers)p.itemsize= 8 bytes (size of a raw 64-bit integer)sys.getsizeof(p[3])= 32 bytes (size of anumpy.int64wrapper object)
Then:
- List calculation:
28 * 1000 = 28000bytes - Correct NumPy calculation:
8 * 1000 = 8000bytes - Your problematic NumPy calculation:
32 * 1000 = 32000bytes
This is why the latter number looks higher—you’re measuring the wrong thing!
Key Takeaway:
NumPy’s memory efficiency comes from storing raw, unboxed data in a contiguous block. To measure its actual memory usage, always use array.itemsize * array.size (or array.nbytes, which does exactly this calculation for you!). Using sys.getsizeof on individual elements will mislead you because those elements are wrapped in Python-compatible objects when accessed directly.
内容的提问来源于stack exchange,提问作者Abhinav Garole

