You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比同名行并新增列标注唯一性原因?

问题:为DataFrame新增Reason列标记同名行的差异列

原始DataFrame

NameAgeCountryOccupationHobby
0A23DEJob holderFishing
1A23DEJob holderGardening
2A23DEJob holderFishing
3A23DEJob holderReading
4B15SWJob holderFishing
5B15SWJob holderPlaying
6C23DDJob holderCoding
7B23AAJob holderFishing
8D34GHJob holderFishing
9D33TROtherFishing

需求

  • 当Name列存在重复值时,对比所有同名行,找出导致该行与其他同名行存在差异的列名,将这些列名用逗号分隔存入新增的Reason列
  • 若某个Name仅出现一次,Reason列填写Unique

预期输出

NameAgeCountryOccupationHobbyReason
0A23DEJob holderFishingOccupation, Hobby
1A23DEJob holderGardeningOccupation, Hobby
2A23DEStudentFishingOccupation, Hobby
3A23DEJob holderReadingOccupation, Hobby
4B15SWJob holderFishingHobby
5B15SWJob holderPlayingHobby
6C23DDJob holderCodingUnique
7B23AAJob holderFishingAge, Country
8D34GHJob holderFishingAge, Country, Occupation
9D33TROtherFishingAge, Country, Occupation

尝试的代码(未得到预期结果)

dif = [i for i, (x,y) in enumerate(zip(df.loc[0].values, df.loc[9,:].values)) if x!=y ]
df.iloc[:, dif]

解决方案

你当前的代码仅对比了第0行和第9行,没有按Name分组处理所有同名行,无法覆盖所有场景。以下是满足需求的实现方案:

完整代码

import pandas as pd

# 构造原始DataFrame
data = [
    ["A",23,"DE","Job holder","Fishing"],
    ["A",23,"DE","Job holder","Gardening"],
    ["A",23,"DE","Student","Fishing"],
    ["A",23,"DE","Job holder","Reading"],
    ["B",15,"SW","Job holder","Fishing"],
    ["B",15,"SW","Job holder","Playing"],
    ["C",23,"DD","Job holder","Coding"],
    ["B",23,"AA","Job holder","Fishing"],
    ["D",34,"GH","Job holder","Fishing"],
    ["D",33,"TR","Other","Fishing"]
]
df = pd.DataFrame(data, columns=["Name","Age","Country","Occupation","Hobby"])

# 定义分组处理函数
def get_reason(group):
    if len(group) == 1:
        return "Unique"
    # 筛选分组内存在不同值的列(排除Name列)
    diff_cols = [col for col in group.columns.drop("Name") if group[col].nunique() > 1]
    return ", ".join(diff_cols)

# 按Name分组处理,将结果映射到原DataFrame
df["Reason"] = df.groupby("Name").transform(get_reason)

print(df)

代码说明

  • groupby("Name"):按Name列分组,确保只处理同名的行
  • transform(get_reason):将分组处理的结果广播到原DataFrame的每一行,保证同一Name的所有行Reason值一致
  • nunique():统计列内不同值的数量,大于1说明该列在分组内存在差异
  • 排除Name列是因为分组依据就是Name,该列值完全相同,无需纳入差异判断

内容的提问来源于stack exchange,提问作者s_max

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 21:00:49