按月份分组计算最高温滚动平均值遇报错,求代码修改方案
问题描述
我用Jupyter、Pandas和Scikit-learn做Python项目的数据预处理,想新增一列monthly_avg,存储每日所在月份的最高温滚动平均值(基于此前所有行计算)。执行以下代码时:
# Group by month in the temperature max column and use the previous rows to calculate the means core_weather["monthly_avg"] = core_weather["temp_max"].groupby(core_weather.index.month).apply(lambda x: x.expanding(1).mean())
出现两个错误:
ValueError: cannot include dtype 'M' in a buffer
TypeError: incompatible index of inserted column with frame index
尝试用core_weather.index = pd.to_datetime(core_weather.index)转换索引类型时,又出现"Only leading negatives allowed"错误。
数据表样本如下:
DATE precip snow snow_depth temp_max temp_min target month_max 1969-06-01 0.00 0.0 0.0 90 72 85.0 NaN 1969-06-02 0.18 0.0 0.0 85 62 80.0 NaN 1969-06-03 0.06 0.0 0.0 80 60 66.0 NaN 1969-06-04 1.27 0.0 0.0 66 60 78.0 NaN 1969-06-05 0.00 0.0 0.0 78 56 86.0 NaN ... ... ... ... ... ... ... ... 2023-09-15 0.02 0.0 0.0 91 71 92.0 100.600000 2023-09-16 0.04 0.0 0.0 92 73 95.0 100.166667 2023-09-17 0.00 0.0 0.0 95 73 94.0 99.900000 2023-09-18 0.00 0.0 0.0 94 69 93.0 99.600000 2023-09-19 0.00 0.0 0.0 93 68 93.0 99.100000
解决方案
1. 修复索引转换问题
从数据样本看,DATE是普通列而非索引,直接转换索引会报错。先将DATE列设为索引,再转换为datetime类型:
# 将DATE列设为索引并转换为datetime格式 core_weather = core_weather.set_index('DATE') core_weather.index = pd.to_datetime(core_weather.index)
如果仍报"Only leading negatives allowed",检查DATE列是否存在格式错误的行(比如非日期格式、负数年份等),清理异常数据后再执行转换。
2. 正确计算月度滚动平均值
原代码用apply会导致索引不匹配,改用groupby+transform可以自动对齐原DataFrame的索引:
# 按月份分组,计算组内从第一行到当前行的滚动平均值 core_weather['monthly_avg'] = core_weather.groupby(core_weather.index.month)['temp_max'].transform(lambda x: x.expanding(1).mean())
关键说明:
groupby(core_weather.index.month):按日期索引的月份(1-12)分组transform:将分组计算的结果映射回原DataFrame,自动保持索引一致性,避免索引不匹配错误expanding(1).mean():计算每个组内从第一行到当前行的累积平均值,即题目要求的"基于此前所有行的滚动均值"
3. 验证结果
执行后可查看前几行数据确认计算是否正确:
print(core_weather[['temp_max', 'monthly_avg']].head())
比如6月1日的monthly_avg应等于当天的temp_max(90),6月2日为(90+85)/2=87.5,以此类推。
内容的提问来源于stack exchange,提问作者ShadowKnight
相关产品推荐
相关产品推荐

