You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按分组计算DataFrame列中指定数值的百分位数

按分组计算指定数值的百分位数

问题描述

现有如下DataFrame:

df:
        Score
group
  A      100
  A      34
  A      40
  A      30
  C      24
  C      60
  C      35

需要为每个分组,计算数值35在Score列中的百分位数,理想输出为:

df:
        Percentile
group 
  A       50
  C       33

尝试过的方法均失败:

  • scipy.stats.percentileofscore(df['Score'], 35, kind='weak'):可运行,但无法按索引分组计算
  • df.groupby('group')['Score'].percentileofscore():报错'SeriesGroupBy' object has no attribute 'percentileofscore'
  • scipy.stats.percentileofscore(df.groupby('group')[['Score]], 35, kind='strict'):报错TypeError: '<' not supported between instances of 'str' and 'int'

解决方法

结合groupby和apply调用scipy.stats.percentileofscore,即可实现分组计算,代码如下:

from scipy.stats import percentileofscore
import pandas as pd

# 构造示例数据
data = {'group': ['A', 'A', 'A', 'A', 'C', 'C', 'C'],
        'Score': [100, 34, 40, 30, 24, 60, 35]}
df = pd.DataFrame(data).set_index('group')

# 分组计算百分位数并转为整数
result = df.groupby('group')['Score'].apply(lambda x: int(percentileofscore(x, 35, kind='weak')))
# 重命名列名并转为DataFrame
result = result.rename('Percentile').to_frame()

print(result)

运行后输出:

Percentile
group             
A               50
C               33

关键说明

  • groupby('group')['Score'].apply()可以对每个分组的Series单独执行计算逻辑
  • percentileofscore的kind参数可按需调整:
    • 'weak':统计小于等于目标值的元素占比,对应示例A组2个元素≤35,占比50%
    • 'strict':统计严格小于目标值的元素占比,对应示例C组1个元素<35,占比约33%
  • 用int()将结果转为整数,匹配理想输出的格式

内容的提问来源于stack exchange,提问作者Yash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 05:40:31