You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中按A、B列去重并保留C列均值

如何用Pandas合并重复行并保留指定列的均值?

给定如下3列数据集:

import pandas as pd

data = {
    'A': [1, 1, 2, 2, 3, 3],
    'B': [11, 11, 7, 12, 10, 10],
    'C': [2.1, 1.4, 2.4, 1.8, 2.6, 2.2]
}
df = pd.DataFrame(data)

数据集展示:

A   B    C  
0  1   11   2.1
1  1   11   1.4
2  2   7    2.4
3  2   12   1.8
4  3   10   2.6
5  3   10   2.2

需求是:将A、B列值相同的重复行合并为一行,同时保留C列的均值。drop_duplicates仅能删除重复行,无法实现聚合计算,因此需要用分组聚合的方式处理。

基础解决方案:使用groupby+mean聚合

通过按A、B列分组,对C列取均值,即可得到核心目标结果:

# 按A、B列分组,计算C列的均值
result = df.groupby(['A', 'B'], as_index=False)['C'].mean()
print(result)

输出结果:

A   B     C
0  1  11  1.75
1  2   7  2.40
2  2  12  1.80
3  3  10  2.40

匹配示例索引的进阶方案

如果需要和需求示例中的行索引完全一致(保留原数据每组的第一个行索引),可以调整代码:

# 添加原索引列,分组时保留每组第一个索引,最后设置为行索引
df['original_index'] = df.index
result = df.groupby(['A', 'B'], as_index=False).agg(
    C=('C', 'mean'),
    original_index=('original_index', 'first')
).set_index('original_index').rename_axis(None)
print(result)

此时输出结果与需求完全匹配:

A   B     C
0  1  11  1.75
2  2   7  2.40
3  2  12  1.80
4  3  10  2.40

内容的提问来源于stack exchange,提问作者Shourov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 07:27:29