基于Pandas实现每日数据按周循环分组处理(规避内存溢出)
Solution for Weekly Batch Processing of Daily Pandas DataFrames
Got it, let's tackle this problem with a memory-efficient approach that avoids loading all your data at once. The core idea is to process one week of data at a time—merge the 7 daily files, aggregate, export, then immediately free up memory before moving to the next week.
Core Logic Breakdown
Instead of loading all 364 days (52 weeks) into memory upfront, we:
- Loop through each week (1 to 52)
- Load only the 7 daily files for that week
- Merge them into a weekly DataFrame
- Add a week identifier, aggregate by
Indvandweek - Export the aggregated data to a file
- Delete unused objects and force garbage collection to free memory
Python Code Implementation
This example assumes your daily data is stored as CSV files named day_1.csv, day_2.csv, etc.—adjust the data loading step to match your actual data source (like Parquet, database queries, etc.):
import pandas as pd import gc # Total weeks to process (52 full weeks = 364 days) total_weeks = 52 for week_num in range(1, total_weeks + 1): # Calculate start/end day indices for the current week start_day = (week_num - 1) * 7 + 1 end_day = week_num * 7 # Collect daily DataFrames for this week daily_data = [] for day in range(start_day, end_day + 1): # Replace this line with your actual data loading logic daily_df = pd.read_csv(f"day_{day}.csv") daily_data.append(daily_df) # Merge daily data into a single weekly DataFrame weekly_df = pd.concat(daily_data, ignore_index=True) # Add a week column to track which week this data belongs to weekly_df["week"] = week_num # Aggregate: sum all numeric columns grouped by Indv and week aggregated_week = weekly_df.groupby(["Indv", "week"]).sum().reset_index() # Export the processed week data (adjust format/path as needed) aggregated_week.to_csv(f"processed_week_{week_num}.csv", index=False) # Clean up memory to prevent buildup del daily_data, weekly_df, aggregated_week gc.collect() print(f"Completed processing Week {week_num}")
Key Customization Tips
- Data Source Adjustment: If your daily data isn't in CSV files, replace
pd.read_csv()with the appropriate method (e.g.,pd.read_parquet()for Parquet files, or a database query usingsqlalchemy). - Custom Aggregations: If you need more than just sums, use
.agg()instead of.sum()to specify different functions for different columns. For example:aggregated_week = weekly_df.groupby(["Indv", "week"]).agg( total_amount=("amount", "sum"), avg_score=("score", "mean") ).reset_index() - Edge Case Handling: If your total days aren't a perfect multiple of 7 (e.g., 365 days), add a final check after the loop to process the remaining days separately.
内容的提问来源于stack exchange,提问作者Srikanth Ayithy
相关产品推荐
相关产品推荐

