如何在DataFrame中按列的不同取值分组计算均值?
按分组计算DataFrame中不同取值对应数据的均值
针对你提到的场景('color'列仅为red/blue),有两种简单的实现方式:
方法一:用groupby()批量分组计算
这是最通用的方法,适合分组后计算单列或多列的均值:
1. 计算指定数值列的分组均值
假设你的DataFrame包含'color'列和需要计算均值的数值列(比如'value'),代码如下:
import pandas as pd # 示例数据 df = pd.DataFrame({ 'color': ['red', 'blue', 'red', 'blue', 'red'], 'value': [10, 20, 15, 25, 12] }) # 按color分组,计算value列的均值 color_mean = df.groupby('color')['value'].mean() print(color_mean)
输出结果:
color blue 22.5 red 12.333333 Name: value, dtype: float64
2. 计算所有数值列的分组均值
如果DataFrame有多列数值型数据,想一次性计算所有列的分组均值,去掉列名索引即可:
all_col_mean = df.groupby('color').mean() print(all_col_mean)
方法二:布尔索引单独计算
因为你的'color'只有red和blue两种取值,也可以直接用布尔筛选后计算均值:
# 计算red对应数据的均值 red_mean = df[df['color'] == 'red']['value'].mean() # 计算blue对应数据的均值 blue_mean = df[df['color'] == 'blue']['value'].mean() print(f"red的均值: {red_mean}, blue的均值: {blue_mean}")
内容的提问来源于stack exchange,提问作者counterculture
相关产品推荐
相关产品推荐

