You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas删除数据行后内存占用升高的原因咨询

Great question! The unexpected increase in memory usage after filtering out NaN rows boils down to a change in the index type between your original customers DataFrame and the filtered customers2 one. Let me break this down step by step:

1. Index Type Difference

  • Your original customers DataFrame uses a RangeIndex (pandas' default for sequentially indexed data). This index is extremely memory-efficient because it only stores three values: the start, stop, and step of the index range. For 378k rows, its memory footprint is negligible (just a few dozen bytes).
  • After filtering with customers[customers['X'].notna() & customers['Y'].notna()], the resulting customers2 uses an Int64Index. This index stores every single index value explicitly (since filtered rows no longer form a continuous range). For 375k rows, each index entry is an 8-byte integer—this adds up to ~2.9 MB of extra memory alone.

2. Memory Breakdown Comparison

Let’s do a rough calculation to see how this adds up:

  • Original customers memory (~8.7 MB):
    • key object column: ~378195 * average UUID string size (~7 bytes) → ~2.6 MB
    • X and Y float64 columns: 2 * 378195 * 8 bytes → ~5.8 MB
    • RangeIndex: ~0 MB
  • Filtered customers2 memory (~11.5 MB):
    • key object column: ~375850 * 7 bytes → ~2.6 MB (slightly reduced)
    • X and Y float64 columns: 2 * 375850 * 8 bytes → ~5.8 MB (slightly reduced)
    • Int64Index: ~2.9 MB (the major new memory cost)
    • Small overhead from non-continuous index structure

3. How to Reduce the Memory Usage

If you want to get back to a leaner memory footprint, reset the index of customers2 to a new RangeIndex:

customers2 = customers2.reset_index(drop=True)

Running customers2.info() after this will show memory usage closer to your expected value, as we’re back to the efficient RangeIndex.

4. Verify the Index Memory Directly

You can confirm this is the root cause by checking the index memory usage explicitly:

# Check original index memory
print(customers.index.memory_usage())
# Check filtered index memory
print(customers2.index.memory_usage())

You’ll see a massive difference between the two values, proving the index type change is the culprit.

内容的提问来源于stack exchange,提问作者Hassen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:08:03