如何用Pandas按跨年度自定义季节分组统计首次末次降雪日?
自定义跨年度季节降雪日期批量统计问题
我拥有1970年至2015年的每日总降雪量数据集,已将日期列设置为索引(格式为year-month-day,如2006-04-24)。需要统计自定义跨年度季节的首次和末次降雪日期,自定义季节为跨年度周期,例如2000/2001季节指2000年6月1日至2001年5月30日。
目前我可通过手动选择时间段,获取该时段内的首次和末次降雪日,代码如下:
df_s = df["2006-04-04" : "2006-04-15"]
firstsnow = df_c[df_c['Height'] > 0].head(1) lastsnow = df_c[df_c['Height'] > 0].tail(1)
但我希望对整个数据集批量处理,按自定义季节分组统计,以便对比各季节降雪日期的变化规律。我推测需使用groupby函数,但不知如何应用。
我的DataFrame结构如下(展示了一段选定时段的数据):Height为降雪高度,Diff为与前一日的差值,二者均为Float64类型:
Height Diff Date 2006-04-04 0.000 NaN 2006-04-05 0.000 0.000 2006-04-06 0.000 0.000 2006-04-07 16.000 16.000 2006-04-08 6.000 -10.000 2006-04-09 0.001 -5.999 2006-04-10 0.000 -0.001 2006-04-11 0.000 0.000 2006-04-12 0.000 0.000 2006-04-13 0.000 0.000 2006-04-14 0.000 0.000 2006-04-15 0.000 0.000
数据类型:<class 'pandas.core.frame.DataFrame'>,形状:(12, 2)
解决方法
核心逻辑是先给每条数据标记所属的自定义季节,再通过groupby按季节分组统计。
1. 生成季节分组标签
自定义季节为6月1日至次年5月30日,需根据日期的月份判断所属季节:
- 月份≥6时,季节为
当年/次年(如2000年6月属于2000/2001季节) - 月份<6时,季节为
上年/当年(如2001年5月属于2000/2001季节)
代码实现:
import pandas as pd # 从索引中提取年份和月份 df['year'] = df.index.year df['month'] = df.index.month # 生成季节标签 df['season'] = df.apply( lambda row: f"{row['year']}/{row['year']+1}" if row['month'] >= 6 else f"{row['year']-1}/{row['year']}", axis=1 )
2. 按季节分组统计首次/末次降雪日期
先筛选出降雪量大于0的记录,再按season分组,取每组的第一个和最后一个日期:
# 筛选有降雪的记录 snow_records = df[df['Height'] > 0].copy() # 分组聚合首次、末次降雪日期 season_snow_stats = snow_records.groupby('season').agg( first_snow_date=('Date', 'first'), last_snow_date=('Date', 'last') )
若你的DataFrame直接用日期作为索引,也可以通过索引聚合:
season_snow_stats = snow_records.groupby('season').agg( first_snow_date=(lambda x: x.index[0]), last_snow_date=(lambda x: x.index[-1]) )
3. 查看结果
执行后season_snow_stats会生成一个以季节为索引的DataFrame,包含每个季节的首次和末次降雪日期,可直接用于后续的规律对比分析。
内容的提问来源于stack exchange,提问作者tryingnottodiewhilecoding
相关产品推荐
相关产品推荐

