如何将三年日粒度数据集的每月数据存储到独立变量/列表?
Got it, let's break this down. You've already got your daily soft drink sales data aggregated in SDD2, and you want to split it into monthly chunks stored separately. Here are a few practical approaches, with the most recommended one first:
1. Use a Dictionary (Highly Recommended)
Creating individual variables for every month can get messy fast—especially with 3 years of data (that's 36 variables!). A dictionary is way cleaner, with month labels as keys and the corresponding monthly data as values.
# First, make sure your Date column is the DatetimeIndex (if it isn't already) # Uncomment the line below if your index isn't already set to Date # SDD2 = SDD2.set_index('Date') # Or ensure the index is datetime type: # SDD2.index = pd.to_datetime(SDD2.index) # Group by month and populate the dictionary monthly_soft_drinks = {} for month, monthly_data in SDD2.groupby(pd.Grouper(freq='M')): # Format the month as a readable key (e.g., "2022-03" for March 2022) month_key = month.strftime('%Y-%m') monthly_soft_drinks[month_key] = monthly_data # Example: Access data for January 2022 print(monthly_soft_drinks['2022-01'])
This way, you can easily iterate through all months, search for specific ones, or modify data without juggling dozens of variable names.
2. Create Individual Variables (Not Recommended, But Possible)
If you really need separate variables (e.g., sdd_202201 for Jan 2022), you can use Python's globals() function to generate them dynamically. Just keep in mind this can clutter your namespace and make debugging harder.
for month, monthly_data in SDD2.groupby(pd.Grouper(freq='M')): # Create a variable name like "sdd_202201" var_name = f"sdd_{month.strftime('%Y%m')}" # Assign the monthly data to the new variable globals()[var_name] = monthly_data # Example: Access the new variable directly print(sdd_202201)
3. Store in a List (For Sequential Access)
If you just need to access months in chronological order without labels, a list works great:
# Group by month and convert to a list monthly_list = [data for _, data in SDD2.groupby(pd.Grouper(freq='M'))] # Example: Access the first month in your dataset print(monthly_list[0]) # Access the 10th month print(monthly_list[9])
Quick Notes:
- The
freq='M'parameter groups data by the last day of each month. If you prefer grouping by the first day, usefreq='MS'instead. - Double-check that your
Dateindex is a datetime type—if not, usepd.to_datetime()to convert it first, otherwise the grouping won't work as expected.
内容的提问来源于stack exchange,提问作者Memon

