使用Pandas删除数据行后内存占用升高的原因咨询
Great question! The unexpected increase in memory usage after filtering out NaN rows boils down to a change in the index type between your original customers DataFrame and the filtered customers2 one. Let me break this down step by step:
1. Index Type Difference
- Your original
customersDataFrame uses aRangeIndex(pandas' default for sequentially indexed data). This index is extremely memory-efficient because it only stores three values: the start, stop, and step of the index range. For 378k rows, its memory footprint is negligible (just a few dozen bytes). - After filtering with
customers[customers['X'].notna() & customers['Y'].notna()], the resultingcustomers2uses anInt64Index. This index stores every single index value explicitly (since filtered rows no longer form a continuous range). For 375k rows, each index entry is an 8-byte integer—this adds up to ~2.9 MB of extra memory alone.
2. Memory Breakdown Comparison
Let’s do a rough calculation to see how this adds up:
- Original
customersmemory (~8.7 MB):keyobject column: ~378195 * average UUID string size (~7 bytes) → ~2.6 MBXandYfloat64 columns: 2 * 378195 * 8 bytes → ~5.8 MB- RangeIndex: ~0 MB
- Filtered
customers2memory (~11.5 MB):keyobject column: ~375850 * 7 bytes → ~2.6 MB (slightly reduced)XandYfloat64 columns: 2 * 375850 * 8 bytes → ~5.8 MB (slightly reduced)- Int64Index: ~2.9 MB (the major new memory cost)
- Small overhead from non-continuous index structure
3. How to Reduce the Memory Usage
If you want to get back to a leaner memory footprint, reset the index of customers2 to a new RangeIndex:
customers2 = customers2.reset_index(drop=True)
Running customers2.info() after this will show memory usage closer to your expected value, as we’re back to the efficient RangeIndex.
4. Verify the Index Memory Directly
You can confirm this is the root cause by checking the index memory usage explicitly:
# Check original index memory print(customers.index.memory_usage()) # Check filtered index memory print(customers2.index.memory_usage())
You’ll see a massive difference between the two values, proving the index type change is the culprit.
内容的提问来源于stack exchange,提问作者Hassen

