You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为DataFrame按五年周期及多维度分组计算指定字段均值?

Hey Guillermo, I’ve got you covered on this one! Let’s walk through how to create that new DataFrame with the five-year period averages you need.

Step 1: Add the Five-Year Period Identifier Column

First, we need to make sure your original table DataFrame has a column that labels each row with its corresponding five-year period.

If you already have your five-year period vector ready, just assign it directly to a new column:

# Replace 'your_five_year_vector' with your actual vector/series
table['FiveYearPeriod'] = your_five_year_vector

If you don’t have the vector mapped yet (e.g., you only have a Year column in your data), you can generate the period labels with a simple function or pd.cut:

# Option 1: Custom function to create human-readable period strings (e.g., "1990-1994")
def assign_five_year_period(year):
    period_start = (year // 5) * 5
    period_end = period_start + 4
    return f"{period_start}-{period_end}"

table['FiveYearPeriod'] = table['Year'].apply(assign_five_year_period)

# Option 2: Use pd.cut for interval-based labels (e.g., "[1990, 1995)")
# Adjust the start/end values to match your data's time range
table['FiveYearPeriod'] = pd.cut(
    table['Year'],
    bins=range(1990, 2021, 5),  # Example: 1990-1994, 1995-1999, ..., 2015-2019
    right=False
).astype(str)  # Convert interval to string for readability

Step 2: Group and Calculate the Averages

Now we’ll group your data by the required columns plus the new FiveYearPeriod column, then compute the mean for tas_month and pr_month:

new_table = table.groupby(
    ['Month', 'Longitude', 'Latitude', 'Model', 'Scenario', 'FiveYearPeriod'],
    as_index=False  # Keeps group columns as regular columns instead of index
)[['tas_month', 'pr_month']].mean()

Quick Notes

  • The as_index=False parameter is key here—it ensures your result is a flat DataFrame instead of a MultiIndex table, which is probably what you want for further analysis.
  • If your data has missing values, you can explicitly skip them in the mean calculation with .mean(skipna=True) (this is the default behavior, but it’s good to be explicit if you want to document it).
  • Double-check that all your grouping columns are in the correct format (e.g., Month is an integer or categorical, not a string like "Jan" unless you want to group by month name).

内容的提问来源于stack exchange,提问作者Guillermo.D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:20:23