You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中defaultdict内存占用过高,32万行CSV处理优化求助

内存优化方案:处理大CSV并按日期分组存储到S3

针对你32GB内存服务器处理32万行CSV时内存耗尽的问题,以下是无需数据库的优化方案,核心思路是避免全量数据驻留内存,改用分块/逐行处理+本地临时文件暂存分组数据:

优化方向1:Pandas分块读取+本地临时文件暂存

利用Pandas的chunksize参数分批次读取CSV,每次仅加载部分数据到内存;按日期分组后将数据追加到对应本地临时文件,最后统一上传到S3。这种方法兼顾Pandas的便捷性和内存效率。

import pandas as pd
import json
from django.core.serializers.json import DjangoJSONEncoder
import boto3
from pathlib import Path

# 初始化S3和临时目录
s3 = boto3.resource('s3')
bucket_storage = '你的存储桶名称'
temp_dir = Path('./temp_date_data')
temp_dir.mkdir(exist_ok=True)

# 分块读取CSV,每次处理10000行(可根据内存调整)
chunk_size = 10000
for chunk in pd.read_csv('目标CSV文件路径.csv', chunksize=chunk_size):
    # 按date分组处理当前块数据
    grouped = chunk.groupby('date')
    for date, group in grouped:
        # 提取location和value并转为字典列表
        data_list = group[['location', 'value']].to_dict('records')
        temp_file_path = temp_dir / f"{date}.json.tmp"
        
        # 合并现有数据(如果临时文件已存在)
        if temp_file_path.exists():
            with open(temp_file_path, 'r') as f:
                existing_data = json.load(f)
            data_list = existing_data + data_list
        
        # 写入临时文件
        with open(temp_file_path, 'w') as f:
            json.dump(data_list, f, cls=DjangoJSONEncoder)

# 所有块处理完成后,上传临时文件到S3并清理
for temp_file in temp_dir.glob('*.json.tmp'):
    date_str = temp_file.stem.replace('.json', '')
    try:
        with open(temp_file, 'rb') as f:
            s3.Object(bucket_storage, f"{date_str}/data.json").put(Body=f)
        temp_file.unlink()  # 删除本地临时文件
    except Exception as e:
        print(f"上传日期{date_str}数据失败: {str(e)}")
temp_dir.rmdir()

优化方向2:用原生CSV模块逐行读取(内存占用最低)

如果Pandas仍占用过多内存,可改用Python原生csv模块逐行读取数据,仅在内存中暂存单个日期的部分数据,达到阈值后写入临时文件,进一步降低内存消耗。

import csv
import json
from django.core.serializers.json import DjangoJSONEncoder
import boto3
from pathlib import Path

# 初始化S3和临时目录
s3 = boto3.resource('s3')
bucket_storage = '你的存储桶名称'
temp_dir = Path('./temp_date_data')
temp_dir.mkdir(exist_ok=True)

# 内存暂存阈值:单个日期数据达到1000条时写入文件
batch_size = 1000
current_data = {}

with open('目标CSV文件路径.csv', 'r') as csv_file:
    reader = csv.DictReader(csv_file)
    for row in reader:
        date = row['date']
        data_item = {'location': row['location'], 'value': row['value']}
        
        # 初始化当前日期的暂存列表(若不存在则读取已有临时文件数据)
        if date not in current_data:
            temp_file_path = temp_dir / f"{date}.json.tmp"
            if temp_file_path.exists():
                with open(temp_file_path, 'r') as f:
                    current_data[date] = json.load(f)
            else:
                current_data[date] = []
        
        current_data[date].append(data_item)
        
        # 达到阈值时写入临时文件并释放内存
        if len(current_data[date]) >= batch_size:
            temp_file_path = temp_dir / f"{date}.json.tmp"
            with open(temp_file_path, 'w') as f:
                json.dump(current_data[date], f, cls=DjangoJSONEncoder)
            del current_data[date]

# 处理剩余未写入的暂存数据
for date, data_list in current_data.items():
    temp_file_path = temp_dir / f"{date}.json.tmp"
    if temp_file_path.exists():
        with open(temp_file_path, 'r') as f:
            existing_data = json.load(f)
        data_list = existing_data + data_list
    with open(temp_file_path, 'w') as f:
        json.dump(data_list, f, cls=DjangoJSONEncoder)

# 上传到S3并清理本地文件
for temp_file in temp_dir.glob('*.json.tmp'):
    date_str = temp_file.stem.replace('.json', '')
    try:
        with open(temp_file, 'rb') as f:
            s3.Object(bucket_storage, f"{date_str}/data.json").put(Body=f)
        temp_file.unlink()
    except Exception as e:
        print(f"上传日期{date_str}数据失败: {str(e)}")
temp_dir.rmdir()

关键优化点说明

  1. 避免全量加载:分块/逐行读取CSV,内存仅保留当前处理的小部分数据,不会一次性加载32万行数据
  2. 替换内存存储:用本地临时文件替代defaultdict存储全部分组数据,内存占用仅为当前处理批次/单个日期的少量数据
  3. 高效分组处理:Pandas的groupby比iterrows效率更高,原生CSV模块则进一步降低内存开销
  4. 减少S3操作:本地合并完单个日期的所有数据后再上传,避免多次S3请求(S3不支持直接追加对象,必须本地合并后上传)

内容的提问来源于stack exchange,提问作者jackhammer013

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 21:06:23