You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python DataFrame聚合某列匹配的所有记录

如何将重复recordhash对应的id合并为逗号分隔的列表?

先修正你原代码里的小问题:df['recordhash'].duplicated(keep=False)返回的是布尔数组,不能直接调用sort_values,得先筛选出重复行再排序:

dups = df[df['recordhash'].duplicated(keep=False)].sort_values('recordhash')

接下来通过分组聚合就能得到你想要的格式:

# 按recordhash分组,合并id为逗号分隔的字符串
result = dups.groupby('recordhash')['id']\
             .apply(lambda x: ', '.join(map(str, x)))\
             .reset_index(name='matching')\
             .reindex(columns=['matching', 'recordhash'])

代码说明:

  • groupby('recordhash')['id']:以recordhash为分组键,只处理id列
  • apply(lambda x: ', '.join(map(str, x))):把每个分组里的id转成字符串后,用逗号加空格连接成一个字符串
  • reset_index(name='matching'):把分组后的索引(recordhash)转为普通列,同时给合并后的id列命名为matching
  • reindex(columns=['matching', 'recordhash']):调整列的顺序,让matching列排在前面

执行后输出就是你想要的格式:

matching recordhash
0    1, 10       ab15

内容的提问来源于stack exchange,提问作者JJ Ryan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 15:15:05