分组后拆分大型Pandas DataFrame及内存优化问题咨询
The problem with your current code is that it loads the entire large CSV into memory upfront, which causes high memory usage. Instead, we can use chunked reading combined with incremental grouping to keep memory usage low while still handling the block-based operations you need. Here's how to do it step by step:
1. Read the CSV in Chunks
Instead of loading the full dataset, use Pandas' chunksize parameter to read the file in smaller, manageable pieces. This way, only one chunk is in memory at a time.
import pandas as pd # Adjust chunksize based on your available memory (e.g., 10,000 rows per chunk) chunk_size = 10000 chunk_iterator = pd.read_table('large_file.csv', sep=',', chunksize=chunk_size)
2. Incrementally Group Data by "block"
Use a dictionary to accumulate data for each block across all chunks. For each chunk, we group by block and append the rows to the corresponding entry in the dictionary:
# Dictionary to hold DataFrames for each block block_data = {} for chunk in chunk_iterator: # Group the current chunk by "block" chunk_groups = chunk.groupby("block") # Append each group's data to the dictionary for block_name, group_df in chunk_groups: if block_name in block_data: block_data[block_name] = pd.concat([block_data[block_name], group_df], ignore_index=True) else: block_data[block_name] = group_df.copy()
3. Process or Write Each Block
Now you can either process each block's data individually or write each block to a separate file (which avoids holding all blocks in memory at once if you process and write one by one):
Option A: Write Each Block to a Separate File
# Write each block to its own CSV file for block_name, df in block_data.items(): df.to_csv(f'block_{block_name}.csv', index=False)
Option B: Process Blocks One by One (Even More Memory-Efficient)
If you don't need to keep all blocks in memory, you can write/process each group as you build it, instead of storing all in the dictionary. This minimizes peak memory usage:
from pathlib import Path # Create a directory to store block files (optional but organized) Path('blocks').mkdir(exist_ok=True) for chunk in chunk_iterator: chunk_groups = chunk.groupby("block") for block_name, group_df in chunk_groups: # Append to the block's file (create if it doesn't exist) file_path = f'blocks/block_{block_name}.csv' group_df.to_csv(file_path, mode='a', header=not Path(file_path).exists(), index=False)
Additional Memory Optimizations
- Specify Data Types: Use the
dtypeparameter inread_tableto assign smaller, appropriate data types (e.g.,dtype={'block': 'category', 'numeric_col': 'int32'}) to reduce memory per chunk. - Load Only Needed Columns: Use
usecolsto load only the columns you need (e.g.,usecols=['block', 'col1', 'col2']), which cuts down on memory usage significantly. - Use
low_memory=False: If you're getting dtype warnings, addlow_memory=Falsetoread_tableto force Pandas to infer types more accurately (though it uses a bit more memory per chunk, it avoids type mismatches).
内容的提问来源于stack exchange,提问作者everestial

