You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按MultiIndex名称对行求和?解决性别占比计算报错

问题解决:计算各职业男性占比并排序(无循环实现)

错误原因分析

你遇到的NameError是因为代码里的occupation是未定义变量,且users.groupby("occupation").gender.value_counts()返回的是MultiIndex(层级索引)——外层为职业(occupation)、内层为性别(gender),无法直接用未定义变量索引。

无循环的实现方案

以下是几种简洁的无循环实现方式,直接完成「按职业+性别分组统计、计算男性占比、排序」的需求:

方案1:利用unstack转换索引后计算

# 按职业和性别分组统计人数,将性别转为列(缺失值用0填充)
gender_stats = users.groupby(['occupation', 'gender']).size().unstack(fill_value=0)
# 计算男性占比 = 男性人数 / 该职业总人数
male_ratio = gender_stats['M'] / gender_stats.sum(axis=1)
# 从高到低排序
male_ratio_sorted = male_ratio.sort_values(ascending=False)

方案2:用value_counts(normalize=True)直接获取占比

这是最简洁的方式,normalize=True会直接返回分组内的占比:

# 按职业分组后统计性别占比,提取男性对应的部分
male_ratio = users.groupby('occupation')['gender'].value_counts(normalize=True).loc[:, 'M']
# 从高到低排序
male_ratio_sorted = male_ratio.sort_values(ascending=False)

这里.loc[:, 'M']利用MultiIndex的索引规则:冒号表示选中所有职业,'M'表示选中男性占比,彻底避免了之前的变量未定义问题。

方案3:结合transform计算总人数

# 给每条数据添加所在职业的总人数列
users['total_occupation'] = users.groupby('occupation')['gender'].transform('count')
# 统计各职业男性人数,除以对应职业总人数得到占比
male_ratio = users[users['gender'] == 'M'].groupby('occupation')['gender'].count() / users.groupby('occupation')['total_occupation'].first()
male_ratio_sorted = male_ratio.sort_values(ascending=False)

内容的提问来源于stack exchange,提问作者doraemon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:55:23