You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用groupby后actual_tmin/actual_tmax格式异常,如何修正?

问题描述

我编写了如下Python代码,用于从亚马逊云服务(AWS)读取指定GHCN站点的日气温数据,计算该站点的历史极端气温及1981-2010年的平均气温:

import pandas as pd

def GHCN(station_ID):
    '''This function reads GHCN Daily Data from Amazon Web Services given a GHCN station ID and calculates the 
       all time record high and low and the normal (mean) high and low temperature for the 1981-2010 period'''

    # Read the GHCN Daily Data
    df = pd.read_csv(
         "s3://noaa-ghcn-pds/csv/by_station/" + station_ID + ".csv",
         storage_options={"anon": True},  # passed to `s3fs.S3FileSystem`
         dtype={'Q_FLAG': 'object', 'M_FLAG': 'object'},
         parse_dates=['DATE']
     ).set_index('DATE')

    df_tmax = df.loc[df['ELEMENT'] == 'TMAX']
    df_tmin = df.loc[df['ELEMENT'] == 'TMIN']

    # Calculate the actual low temperature
    ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))]
    actual_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year)

    # Calculate the actual high temperature
    ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))]
    actual_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year)

    # Calculate the record low temperature
    ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))]
    record_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).min()

    # Calculate the record high temperature
    ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))]
    record_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).max()

    # Calculate the mean low temperature
    ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))]
    mean_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).mean()

    # Calculate the mean high temperature
    ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))]
    mean_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).mean()

    # Create pandas dataframe for temperature values
    temps = pd.DataFrame([[actual_tmin, mean_tmin, record_tmin, actual_tmax, mean_tmax, record_tmax]],
        columns= ['actual_min_temp', 'average_min_temp', 'record_min_temp', 'actual_max_temp', 'average_max_temp', 'record_max_temp'])
    
    return temps, actual_tmin, actual_tmax, record_tmin, record_tmax, mean_tmin, mean_tmax

GHCN('USW00094870')

调用该函数后,发现actual_tmin和actual_tmax为<pandas.core.groupby.generic.SeriesGroupBy>对象,格式与record_tmin、mean_tmin等已完成聚合的气温数据不同,输出示例如下:

(                                     actual_min_temp  
 0  <pandas.core.groupby.generic.SeriesGroupBy obj...   
 
                                     average_min_temp  
 0  DATE
 1     -8.076000
 2     -6.860000
 3     -7....   

我想请教如何让actual_tmax和actual_tmin的格式与其他计算出的气温数据一致?

解决方案

问题核心是actual_tmin和actual_tmax只执行了groupby()分组操作,没有调用聚合函数(如.min()/.mean()),而其他变量都通过聚合方法生成了最终的Series结果。根据需求,有两种常见修改方式:

1. 获取每个日序的所有实际观测值

如果需要保留每个日序对应的所有历史气温观测值,用.apply(list)将分组内的值转换成列表,得到与其他变量格式一致的Series:

# 修改actual_tmin的计算
ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))]
actual_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).apply(list)

# 修改actual_tmax的计算
ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))]
actual_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).apply(list)

2. 获取每个日序的最新(最近年份)实际气温值

如果只需要每个日序对应的最新一次观测气温,用.last()聚合方法:

# 修改actual_tmin的计算
ser=df_tmin[~((df_tmin.index.month==2)&(df_tmin.index.day==29))]
actual_tmin = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).last()

# 修改actual_tmax的计算
ser=df_tmax[~((df_tmax.index.month==2)&(df_tmax.index.day==29))]
actual_tmax = (ser['DATA_VALUE']/10.).groupby(ser.index.day_of_year).last()

额外优化建议

代码中重复了大量相同的数据过滤逻辑,可提前处理好过滤后的数据集,减少冗余计算:

import pandas as pd

def GHCN(station_ID):
    '''This function reads GHCN Daily Data from Amazon Web Services given a GHCN station ID and calculates the 
       all time record high and low and the normal (mean) high and low temperature for the 1981-2010 period'''

    # 读取GHCN数据
    df = pd.read_csv(
         "s3://noaa-ghcn-pds/csv/by_station/" + station_ID + ".csv",
         storage_options={"anon": True},
         dtype={'Q_FLAG': 'object', 'M_FLAG': 'object'},
         parse_dates=['DATE']
     ).set_index('DATE')

    # 提前过滤2月29日数据,并拆分TMAX/TMIN,转换温度单位
    filter_leap_day = ~((df.index.month == 2) & (df.index.day == 29))
    df_tmax = df.loc[(df['ELEMENT'] == 'TMAX') & filter_leap_day]['DATA_VALUE'] / 10.
    df_tmin = df.loc[(df['ELEMENT'] == 'TMIN') & filter_leap_day]['DATA_VALUE'] / 10.

    # 按日序分组
    tmax_grouped = df_tmax.groupby(df_tmax.index.day_of_year)
    tmin_grouped = df_tmin.groupby(df_tmin.index.day_of_year)

    # 计算各类气温值(根据需求选择actual的聚合方式)
    actual_tmin = tmin_grouped.apply(list)  # 或tmin_grouped.last()
    actual_tmax = tmax_grouped.apply(list)  # 或tmax_grouped.last()
    record_tmin = tmin_grouped.min()
    record_tmax = tmax_grouped.max()
    mean_tmin = tmin_grouped.mean()
    mean_tmax = tmax_grouped.mean()

    # 构建结果DataFrame
    temps = pd.DataFrame({
        'actual_min_temp': [actual_tmin],
        'average_min_temp': [mean_tmin],
        'record_min_temp': [record_tmin],
        'actual_max_temp': [actual_tmax],
        'average_max_temp': [mean_tmax],
        'record_max_temp': [record_tmax]
    })
    
    return temps, actual_tmin, actual_tmax, record_tmin, record_tmax, mean_tmin, mean_tmax

GHCN('USW00094870')

内容的提问来源于stack exchange,提问作者kjmw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 06:55:21