对比DataFrame子集与全集并计算占比:求解拼写比赛晋级最后轮次儿童的百分比问题
计算晋级最后一轮儿童的占比百分比
嘿,这个需求用Pandas其实很容易实现,我来给你拆解步骤,顺便分享几个实用的小技巧~
首先咱们核心要拿到两个数据:总参赛儿童数和成功晋级最后一轮的儿童数,用后者除以前者再乘以100就能得到百分比。下面分两种常见场景给你具体代码:
场景1:晋级轮次是数值型(比如1、2、3…最后一轮是最大值)
假设你的DataFrame叫spell_comp_df,记录晋级轮次的列名为final_round:
# 1. 获取最后一轮的轮次号(数值型列的最大值就是最后一轮) last_round = spell_comp_df['final_round'].max() # 2. 计算总参赛人数 total_kids = spell_comp_df.shape[0] # 也可以用 len(spell_comp_df),效果完全一致 # 3. 筛选出晋级最后一轮的人数 finalists = spell_comp_df[spell_comp_df['final_round'] == last_round].shape[0] # 4. 计算百分比,保留两位小数更直观 percentage = round((finalists / total_kids) * 100, 2) print(f"晋级最后一轮的儿童占比:{percentage}%")
场景2:晋级轮次是字符串型(比如"Round 1"、"Final Round")
如果轮次列是明确的文字标识,直接匹配目标值就行:
# 直接筛选出晋级"Final Round"的行 finalists = spell_comp_df[spell_comp_df['final_round'] == 'Final Round'].shape[0] total_kids = len(spell_comp_df) percentage = round((finalists / total_kids) * 100, 2) print(f"晋级最后一轮的儿童占比:{percentage}%")
更简洁的技巧:用mean()一步算比例
Pandas里布尔类型的Series平均值就是符合条件的比例(因为True会被当作1,False当作0),所以可以省掉好几行代码:
# 数值型轮次的情况 percentage = round((spell_comp_df['final_round'] == spell_comp_df['final_round'].max()).mean() * 100, 2) # 字符串型轮次的情况 percentage = round((spell_comp_df['final_round'] == 'Final Round').mean() * 100, 2)
常用工具函数说明
df.shape[0]:快速获取DataFrame的行数(总人数),大数据集下比len(df)效率略高df[condition].shape[0]:筛选符合条件的行后获取行数(晋级人数)Series.mean():利用布尔值特性直接计算比例,代码更简洁round():控制百分比的小数位数,让结果更美观
内容的提问来源于stack exchange,提问作者cxspv2108
相关产品推荐
相关产品推荐

