Pandas如何按区域分组求取数据框各列最大值及对应国家
实现方案
方法1:基于已有的区域列表遍历实现
你已经通过np.unique(df['Region'])拿到了区域列表,后续可以按如下逻辑操作:
- 初始化空容器存储最终结果
- 遍历每个区域,筛选出该区域对应的所有数据
- 分别匹配
Happiness score、GDP两列最大值对应的行,提取数值和国家名称 - 汇总所有区域的结果
import pandas as pd import numpy as np # 你已生成的区域列表 regions = np.unique(df['Region']) result = [] for region in regions: # 筛选当前区域的全量数据 region_df = df[df['Region'] == region].reset_index(drop=True) # 取幸福指数最高的对应行 happiness_max_row = region_df.loc[region_df['Happiness score'].idxmax()] # 取GDP最高的对应行 gdp_max_row = region_df.loc[region_df['GDP'].idxmax()] # 存入单区域结果 result.append({ 'Region': region, 'max_happiness_score': happiness_max_row['Happiness score'], 'happiness_max_country': happiness_max_row['country name'], 'max_gdp': gdp_max_row['GDP'], 'gdp_max_country': gdp_max_row['country name'] }) # 转换为数据框方便后续处理 result_df = pd.DataFrame(result)
方法2:groupby聚合实现(更推荐,性能更高)
不需要单独提前提取区域列表,直接用pandas内置的分组聚合能力实现,代码更简洁,处理大数据量时性能远高于遍历:
# 自定义聚合函数:返回最大值+对应国家名 def agg_max_info(series): max_value = series.max() max_country = df.loc[series.idxmax(), 'country name'] return pd.Series([max_value, max_country], index=['max_value', 'corresponding_country']) # 按区域分组,对两个数值列做聚合 result_df = df.groupby('Region')[['Happiness score', 'GDP']].apply(agg_max_info).unstack() # 重命名列名,优化可读性 result_df.columns = [f'{col[0]}_{col[1]}' for col in result_df.columns] result_df = result_df.reset_index()
注意:如果单区域内存在多个国家数值同为最大值的情况,上述方法默认返回第一个匹配到的国家,若需要返回所有匹配国家,可将聚合逻辑中的国家返回值修改为列表存储所有匹配结果。
内容的提问来源于stack exchange,提问作者jim
相关产品推荐
相关产品推荐

