You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python pandas按指定列分组并计算新的百分比字段?

问题解决:pandas按水果分组计算Good状态占比

解决代码

你可以参考下面两种实现方式,都可以得到你需要的输出结果:

写法1:直接基于分组自定义计算

import pandas as pd
df = pd.DataFrame({'Fruit': ['Apple','Apple','Banana'], 'Condition': ['Good','Bad','Good']})

result = df.groupby('Fruit', as_index=False).apply(
    lambda group: pd.Series({
        'Percentage': f"{(group['Condition'] == 'Good').sum() / len(group) * 100:.0f}%"
    })
)
print(result)

写法2:性能更优的向量化实现

适合数据量较大的场景,避免逐组调用lambda的性能损耗:

import pandas as pd
df = pd.DataFrame({'Fruit': ['Apple','Apple','Banana'], 'Condition': ['Good','Bad','Good']})

# 新增临时列标记是否为Good状态
df['is_good'] = df['Condition'] == 'Good'
# 分组计算Good状态占比
result = df.groupby('Fruit', as_index=False)['is_good'].mean()
# 格式化为百分比字符串
result['Percentage'] = result['is_good'].map(lambda x: f"{x*100:.0f}%")
# 清理临时列得到最终结果
result = result.drop(columns=['is_good'])
print(result)

输出结果

两种写法的输出都符合你的预期:

Fruit Percentage
0   Apple        50%
1  Banana       100%

原代码问题原因

你之前的代码得到两列相同的比值,是因为x[(x['Condition'] == 'Good')].count()会统计筛选后DataFrame所有列的非空值数量,再除以原分组所有列的非空值数量,所以Fruit和Condition列都会返回相同的占比结果。你只需要在lambda中指定仅计算Condition列的符合条件占比,并且返回带自定义列名的Series即可解决。

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 12:24:07