You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas扩展DataFrame列内存占用过高,求低内存解决方案

低内存实现范围行扩展的解决方案

针对10亿行规模的扩展需求,核心思路是避免在内存中存储所有扩展后的数据,而是分块处理并直接将结果写入磁盘,同时尽可能压缩数据类型。以下是两种高效方案:

方案一:基于Numpy的批量块处理(平衡速度与内存)

利用Numpy的向量化操作批量生成扩展数据,处理完一个块就写入磁盘并释放内存,避免累积:

import pandas as pd
import numpy as np

chunk_size = 100000  # 保持你的分块大小
output_path = "expanded_result.csv"

# 先写入表头
pd.DataFrame(columns=["FID", "SID"]).to_csv(output_path, index=False)

# 分块读取原数据,强制指定最小可行数据类型
for chunk in pd.read_csv(
    "your_input_data.csv",
    chunksize=chunk_size,
    dtype={"FID": "int32", "SID_START": "int32", "SID_END": "int32"}
):
    # 计算每行需要扩展的行数
    expand_lengths = chunk["SID_END"] - chunk["SID_START"] + 1
    # 批量重复FID
    repeated_fids = np.repeat(chunk["FID"].values, expand_lengths)
    # 批量生成所有SID序列,强制转int32节省内存
    sid_sequences = [
        np.arange(start, end + 1, dtype="int32")
        for start, end in zip(chunk["SID_START"], chunk["SID_END"])
    ]
    combined_sids = np.concatenate(sid_sequences)
    # 构建临时DF并追加到输出文件
    temp_df = pd.DataFrame(
        {"FID": repeated_fids, "SID": combined_sids},
        dtype="int32"
    )
    temp_df.to_csv(output_path, mode="a", header=False, index=False)
    # 手动释放内存,避免累积
    del repeated_fids, sid_sequences, combined_sids, temp_df

优势:

  • 向量化操作比逐行循环更快
  • 每次仅保留一个块的扩展数据,写完即释放,内存占用稳定在分块对应扩展后的规模(远低于155GB)
  • 强制指定int32进一步压缩临时数组的内存占用

方案二:逐行生成写入(极致低内存)

如果内存资源极度紧张,直接逐行处理原数据,生成一行扩展数据就写入一行,内存仅保留当前处理的单行数据:

import pandas as pd

chunk_size = 100000
output_path = "expanded_result.csv"

# 写入表头
with open(output_path, "w", encoding="utf-8") as f:
    f.write("FID,SID\n")

# 分块读取原数据
for chunk in pd.read_csv(
    "your_input_data.csv",
    chunksize=chunk_size,
    dtype={"FID": "int32", "SID_START": "int32", "SID_END": "int32"}
):
    with open(output_path, "a", encoding="utf-8") as f:
        # 逐行遍历块内数据
        for _, row in chunk.iterrows():
            current_fid = row["FID"]
            # 生成当前行的所有SID并写入
            for sid in range(row["SID_START"], row["SID_END"] + 1):
                f.write(f"{current_fid},{sid}\n")
    # 释放块内存
    del chunk

优势:

  • 内存占用极低(仅当前块和单行数据)
  • 无需生成大的中间数组,完全避免内存过载

关键注意事项

  1. 数据类型压缩:确认FID、SID的数值范围,尽可能用最小的整数类型(如int16/int32),不要默认用int64,可直接减少一半以上的内存占用。
  2. 避免内存累积:绝对不要在内存中保存完整的扩展数据集,所有处理结果直接写入磁盘。
  3. 弃用repeat+cumcount:原方法会在内存中生成与扩展后数据规模一致的中间数组,对于大分块来说必然导致内存暴涨,完全不适合超大规模数据。

内容的提问来源于stack exchange,提问作者abcd_1234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 22:02:10