如何使用Python pandas按指定列分组并计算新的百分比字段?
问题解决:pandas按水果分组计算Good状态占比
解决代码
你可以参考下面两种实现方式,都可以得到你需要的输出结果:
写法1:直接基于分组自定义计算
import pandas as pd df = pd.DataFrame({'Fruit': ['Apple','Apple','Banana'], 'Condition': ['Good','Bad','Good']}) result = df.groupby('Fruit', as_index=False).apply( lambda group: pd.Series({ 'Percentage': f"{(group['Condition'] == 'Good').sum() / len(group) * 100:.0f}%" }) ) print(result)
写法2:性能更优的向量化实现
适合数据量较大的场景,避免逐组调用lambda的性能损耗:
import pandas as pd df = pd.DataFrame({'Fruit': ['Apple','Apple','Banana'], 'Condition': ['Good','Bad','Good']}) # 新增临时列标记是否为Good状态 df['is_good'] = df['Condition'] == 'Good' # 分组计算Good状态占比 result = df.groupby('Fruit', as_index=False)['is_good'].mean() # 格式化为百分比字符串 result['Percentage'] = result['is_good'].map(lambda x: f"{x*100:.0f}%") # 清理临时列得到最终结果 result = result.drop(columns=['is_good']) print(result)
输出结果
两种写法的输出都符合你的预期:
Fruit Percentage 0 Apple 50% 1 Banana 100%
原代码问题原因
你之前的代码得到两列相同的比值,是因为x[(x['Condition'] == 'Good')].count()会统计筛选后DataFrame所有列的非空值数量,再除以原分组所有列的非空值数量,所以Fruit和Condition列都会返回相同的占比结果。你只需要在lambda中指定仅计算Condition列的符合条件占比,并且返回带自定义列名的Series即可解决。
内容的提问来源于stack exchange,提问作者Chris
相关产品推荐
相关产品推荐

