如何在Python中按年份计算每日DataFrame的月度均值?
Hey there! The issue with your original code is that you're grouping only by month name—so it's aggregating all January data from every year together, all February data together, etc. That's why you're getting total sums per month across all years instead of year-by-year monthly averages.
Let's walk through two straightforward ways to get the year-by-year monthly mean you're looking for:
Method 1: Using resample()
First, make sure your ds column is formatted as a datetime type (if it isn't already):
results['ds'] = pd.to_datetime(results['ds'])
Then use resample() with the 'M' rule (for monthly frequency) and specify the on parameter to tell pandas which datetime column to use. We'll use .mean() to calculate the average for each month:
# Get year-by-year monthly averages yearly_monthly_mean = results.resample('M', on='ds')['y'].mean()
By default, the resulting index will be the last day of each month (e.g., 2000-01-31). If you prefer the index to show the first day of the month instead, add the label='left' parameter:
yearly_monthly_mean = results.resample('M', on='ds', label='left')['y'].mean()
The output will look something like this (using your sample data):
ds 2000-01-01 1.828917 ... 2018-04-01 3.426111 Name: y, dtype: float64
Method 2: Using groupby() with year and month
If you want more control over how the groups are labeled (e.g., showing month names instead of numbers), you can group by both year and month:
Option A: Group by year and month number
yearly_monthly_mean = results.groupby([ results['ds'].dt.year, results['ds'].dt.month ])['y'].mean()
This will give you a multi-index where the first level is the year, and the second is the month number.
Option B: Group by year and month name (for readability)
If you want month names instead of numbers, use dt.strftime('%B'), and optionally unstack() to pivot the months into columns for a cleaner view:
yearly_monthly_mean = results.groupby([ results['ds'].dt.year, results['ds'].dt.strftime('%B') ])['y'].mean().unstack()
This will produce a DataFrame with years as rows and month names as columns, making it easy to compare monthly averages across years.
Why your original code didn't work
Your line results.groupby(results['ds'].dt.strftime('%B'))['y'].sum().sort_values() only groups by month name (e.g., "January", "February")—so all data from every January in your dataset gets summed together. Adding the year to the grouping key ensures you get separate aggregates for each year-month combination.
内容的提问来源于stack exchange,提问作者Cesar

