You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按rev_id分组求toxicity众数与toxicity_score均值?

如何用Pandas按rev_id分组,获取toxicity的众数和toxicity_score的均值?

没问题,这是用Pandas实现你需求的具体方法:

完整代码示例

首先,先把你提供的数据集转换成Pandas DataFrame:

import pandas as pd

# 构造你的数据集
data = {
    'rev_id': [2232.0]*10 + [4216.0]*10 + [8953.0]*10,
    'worker_id': [723,4000,3989,3341,1574,1508,772,680,405,4020,
                  500,599,339,257,303,188,1549,64,1527,1502,
                  2596,2403,2539,2542,2544,1016,2550,2578,2494,971],
    'toxicity': [0]*9 + [1] + [0]*10 + [0]*10,
    'toxicity_score': [0.0,0.0,1.0,0.0,1.0,1.0,1.0,0.0,1.0,-1.0,
                       0.0,0.0,2.0,0.0,1.0,0.0,1.0,1.0,0.0,0.0,
                       1.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0,0.0]
}

df = pd.DataFrame(data)

然后执行分组聚合操作:

# 按rev_id分组,聚合toxicity的众数和toxicity_score的均值
result = df.groupby('rev_id').agg(
    toxicity_mode=('toxicity', lambda x: x.mode().iloc[0]),
    toxicity_score_mean=('toxicity_score', 'mean')
).reset_index()

print(result)

代码解释

  • groupby('rev_id'):将数据按照rev_id字段进行分组
  • agg():指定每个字段的聚合规则:
    • 对于toxicity字段,我们用lambda x: x.mode().iloc[0]来获取众数:mode()方法会返回该分组中出现次数最多的值(如果有多个众数,会返回所有值,这里因为toxicity只有0和1,我们取第一个结果即可)
    • 对于toxicity_score字段,直接用'mean'即可计算该分组的均值
  • reset_index():将分组后的索引(rev_id)转换为普通列,让结果更直观

输出结果

运行代码后,你会得到如下结果:

rev_id  toxicity_mode  toxicity_score_mean
0  2232.0              0                  0.4
1  4216.0              0                  0.5
2  8953.0              0                  0.1

内容的提问来源于stack exchange,提问作者yanachen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:00:37