You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分组后拆分大型Pandas DataFrame及内存优化问题咨询

Solution to Memory-Efficient Grouped Processing of Large CSV Files with Pandas

The problem with your current code is that it loads the entire large CSV into memory upfront, which causes high memory usage. Instead, we can use chunked reading combined with incremental grouping to keep memory usage low while still handling the block-based operations you need. Here's how to do it step by step:

1. Read the CSV in Chunks

Instead of loading the full dataset, use Pandas' chunksize parameter to read the file in smaller, manageable pieces. This way, only one chunk is in memory at a time.

import pandas as pd

# Adjust chunksize based on your available memory (e.g., 10,000 rows per chunk)
chunk_size = 10000
chunk_iterator = pd.read_table('large_file.csv', sep=',', chunksize=chunk_size)

2. Incrementally Group Data by "block"

Use a dictionary to accumulate data for each block across all chunks. For each chunk, we group by block and append the rows to the corresponding entry in the dictionary:

# Dictionary to hold DataFrames for each block
block_data = {}

for chunk in chunk_iterator:
    # Group the current chunk by "block"
    chunk_groups = chunk.groupby("block")
    
    # Append each group's data to the dictionary
    for block_name, group_df in chunk_groups:
        if block_name in block_data:
            block_data[block_name] = pd.concat([block_data[block_name], group_df], ignore_index=True)
        else:
            block_data[block_name] = group_df.copy()

3. Process or Write Each Block

Now you can either process each block's data individually or write each block to a separate file (which avoids holding all blocks in memory at once if you process and write one by one):

Option A: Write Each Block to a Separate File

# Write each block to its own CSV file
for block_name, df in block_data.items():
    df.to_csv(f'block_{block_name}.csv', index=False)

Option B: Process Blocks One by One (Even More Memory-Efficient)

If you don't need to keep all blocks in memory, you can write/process each group as you build it, instead of storing all in the dictionary. This minimizes peak memory usage:

from pathlib import Path

# Create a directory to store block files (optional but organized)
Path('blocks').mkdir(exist_ok=True)

for chunk in chunk_iterator:
    chunk_groups = chunk.groupby("block")
    for block_name, group_df in chunk_groups:
        # Append to the block's file (create if it doesn't exist)
        file_path = f'blocks/block_{block_name}.csv'
        group_df.to_csv(file_path, mode='a', header=not Path(file_path).exists(), index=False)

Additional Memory Optimizations

  • Specify Data Types: Use the dtype parameter in read_table to assign smaller, appropriate data types (e.g., dtype={'block': 'category', 'numeric_col': 'int32'}) to reduce memory per chunk.
  • Load Only Needed Columns: Use usecols to load only the columns you need (e.g., usecols=['block', 'col1', 'col2']), which cuts down on memory usage significantly.
  • Use low_memory=False: If you're getting dtype warnings, add low_memory=False to read_table to force Pandas to infer types more accurately (though it uses a bit more memory per chunk, it avoids type mismatches).

内容的提问来源于stack exchange,提问作者everestial

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:43:38