Python Groupby结合排序:求各职业男性占比并降序排序遇阻求助
计算各职业男性占比并排序的完整解决方案
嘿,我来帮你搞定这个问题!计算各职业男性占比并排序其实很清晰,我给你分步拆解,附上两种常用的实现方法,你可以根据自己的DataFrame结构灵活调整:
前提假设
先假设你的DataFrame包含两列核心数据:
occupation:职业名称列gender:性别列(取值比如'M'代表男性,'F'代表女性;如果你的数据是数字编码,比如1/0,后面的代码只需微调判断条件即可)
如果你的DataFrame有缺失值,建议先处理:
# 移除职业或性别列的缺失行 df = df.dropna(subset=['occupation', 'gender'])
方法一:Groupby + 聚合计算
这是最直观的分步写法,适合新手理解每一步逻辑:
import pandas as pd # 1. 按职业分组,统计总人数和男性人数 occupation_stats = df.groupby('occupation').agg( total_people=('gender', 'count'), # 该职业总人数 male_people=('gender', lambda x: (x == 'M').sum()) # 该职业男性人数 ) # 2. 计算男性占比(保留2位小数,便于阅读) occupation_stats['male_ratio'] = (occupation_stats['male_people'] / occupation_stats['total_people']).round(2) # 3. 按男性占比从高到低排序 sorted_result = occupation_stats.sort_values(by='male_ratio', ascending=False) # 查看最终结果 print(sorted_result)
方法二:Pivot_table 简洁写法
如果习惯更紧凑的代码,用透视表可以一步完成性别分布统计:
import pandas as pd # 1. 生成各职业的性别分布透视表 gender_pivot = pd.pivot_table( df, index='occupation', columns='gender', aggfunc='size', # 统计数量 fill_value=0 # 把空值填充为0 ) # 2. 计算男性占比 gender_pivot['male_ratio'] = (gender_pivot['M'] / (gender_pivot['M'] + gender_pivot['F'])).round(2) # 3. 按占比降序排序 sorted_pivot = gender_pivot.sort_values(by='male_ratio', ascending=False) print(sorted_pivot)
小提示
- 如果你的性别列是其他取值(比如
1代表男,0代表女),只需把代码里的x == 'M'改成x == 1即可 - 要是想把占比转成百分比格式(比如
85%),可以用:occupation_stats['male_ratio'] = (occupation_stats['male_ratio'] * 100).astype(str) + '%'
内容的提问来源于stack exchange,提问作者subodh agrawal
相关产品推荐
相关产品推荐

