使用groupby后actual_tmin/actual_tmax格式异常,如何修正?
问题描述
我编写了如下Python代码,用于从亚马逊云服务(AWS)读取指定GHCN站点的日气温数据,计算该站点的历史极端气温及1981-2010年的平均气温:
import pandas as pd def GHCN(station_ID): '''This function reads GHCN Daily Data from Amazon Web Services given a GHCN station ID and calculates the all time record high and low and the normal (mean) high and low temperature for the 1981-2010 period''' # Read the GHCN Daily Data df = pd.read_csv( "s3://noaa-ghcn-pds/csv/by_station/" + station_ID + ".csv", storage_options={"anon": True}, # passed to `s3fs.S3FileSystem` dtype={'Q_FLAG': 'object', 'M_FLAG': 'object'}, parse_dates=['DATE'] ).set_index('DATE') df_tmax = df.loc[df['ELEMENT'] == 'TMAX'] df_tmin = df.loc[df['ELEMENT'] == 'TMIN'] # Calculate the actual low temperature ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))] actual_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year) # Calculate the actual high temperature ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))] actual_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year) # Calculate the record low temperature ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))] record_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).min() # Calculate the record high temperature ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))] record_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).max() # Calculate the mean low temperature ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))] mean_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).mean() # Calculate the mean high temperature ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))] mean_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).mean() # Create pandas dataframe for temperature values temps = pd.DataFrame([[actual_tmin, mean_tmin, record_tmin, actual_tmax, mean_tmax, record_tmax]], columns= ['actual_min_temp', 'average_min_temp', 'record_min_temp', 'actual_max_temp', 'average_max_temp', 'record_max_temp']) return temps, actual_tmin, actual_tmax, record_tmin, record_tmax, mean_tmin, mean_tmax GHCN('USW00094870')
调用该函数后,发现actual_tmin和actual_tmax为<pandas.core.groupby.generic.SeriesGroupBy>对象,格式与record_tmin、mean_tmin等已完成聚合的气温数据不同,输出示例如下:
( actual_min_temp 0 <pandas.core.groupby.generic.SeriesGroupBy obj... average_min_temp 0 DATE 1 -8.076000 2 -6.860000 3 -7....
我想请教如何让actual_tmax和actual_tmin的格式与其他计算出的气温数据一致?
解决方案
问题核心是actual_tmin和actual_tmax只执行了groupby()分组操作,没有调用聚合函数(如.min()/.mean()),而其他变量都通过聚合方法生成了最终的Series结果。根据需求,有两种常见修改方式:
1. 获取每个日序的所有实际观测值
如果需要保留每个日序对应的所有历史气温观测值,用.apply(list)将分组内的值转换成列表,得到与其他变量格式一致的Series:
# 修改actual_tmin的计算 ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))] actual_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).apply(list) # 修改actual_tmax的计算 ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))] actual_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).apply(list)
2. 获取每个日序的最新(最近年份)实际气温值
如果只需要每个日序对应的最新一次观测气温,用.last()聚合方法:
# 修改actual_tmin的计算 ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))] actual_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).last() # 修改actual_tmax的计算 ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))] actual_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).last()
额外优化建议
代码中重复了大量相同的数据过滤逻辑,可提前处理好过滤后的数据集,减少冗余计算:
import pandas as pd def GHCN(station_ID): '''This function reads GHCN Daily Data from Amazon Web Services given a GHCN station ID and calculates the all time record high and low and the normal (mean) high and low temperature for the 1981-2010 period''' # 读取GHCN数据 df = pd.read_csv( "s3://noaa-ghcn-pds/csv/by_station/" + station_ID + ".csv", storage_options={"anon": True}, dtype={'Q_FLAG': 'object', 'M_FLAG': 'object'}, parse_dates=['DATE'] ).set_index('DATE') # 提前过滤2月29日数据,并拆分TMAX/TMIN,转换温度单位 filter_leap_day = ~((df.index.month == 2) & (df.index.day == 29)) df_tmax = df.loc[(df['ELEMENT'] == 'TMAX') & filter_leap_day]['DATA_VALUE'] / 10. df_tmin = df.loc[(df['ELEMENT'] == 'TMIN') & filter_leap_day]['DATA_VALUE'] / 10. # 按日序分组 tmax_grouped = df_tmax.groupby(df_tmax.index.day_of_year) tmin_grouped = df_tmin.groupby(df_tmin.index.day_of_year) # 计算各类气温值(根据需求选择actual的聚合方式) actual_tmin = tmin_grouped.apply(list) # 或tmin_grouped.last() actual_tmax = tmax_grouped.apply(list) # 或tmax_grouped.last() record_tmin = tmin_grouped.min() record_tmax = tmax_grouped.max() mean_tmin = tmin_grouped.mean() mean_tmax = tmax_grouped.mean() # 构建结果DataFrame temps = pd.DataFrame({ 'actual_min_temp': [actual_tmin], 'average_min_temp': [mean_tmin], 'record_min_temp': [record_tmin], 'actual_max_temp': [actual_tmax], 'average_max_temp': [mean_tmax], 'record_max_temp': [record_tmax] }) return temps, actual_tmin, actual_tmax, record_tmin, record_tmax, mean_tmin, mean_tmax GHCN('USW00094870')
内容的提问来源于stack exchange,提问作者kjmw
相关产品推荐
相关产品推荐

