You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现Pandas中任意数据类型的两行差异对比?函数报错求助

解决Pandas两行任意数据类型对比的ValueError问题

错误原因

你写的代码里,df.loc[df['Employee Name'] == row1][col1]返回的是带原行索引的Series(比如John对应索引0,Jill对应索引1),两个索引不同的Series直接比较时,Pandas会要求必须是"identically-labeled"(索引完全一致)的对象,因此触发ValueError。

修正方案

方法1:提取单行标量值逐列对比

先把目标行转成字典,直接按列取值对比:

def sim_func(df, row1, row2):
    # 提取两行的标量数据转成字典
    row_data1 = df[df['Employee Name'] == row1].iloc[0].to_dict()
    row_data2 = df[df['Employee Name'] == row2].iloc[0].to_dict()
    
    # 遍历列名,输出值不同的列
    for col in df.columns:
        if row_data1[col] != row_data2[col]:
            print(col)

sim_func(cars, 'John', 'Jill')

运行输出:

Sex
# Cars Sold
Years With Company
Personal Car

方法2:Series重置索引后对比

提取两行的Series并重置索引,确保索引一致后直接对比:

def sim_func(df, row1, row2):
    # 获取两行的Series,重置索引消除原行号影响
    s1 = df[df['Employee Name'] == row1].iloc[0].reset_index(drop=True)
    s2 = df[df['Employee Name'] == row2].iloc[0].reset_index(drop=True)
    
    # 筛选并输出值不同的列名
    diff_cols = s1[s1 != s2].index.tolist()
    for col in diff_cols:
        print(col)

sim_func(cars, 'John', 'Jill')

方法3:一行代码快速实现

直接用布尔筛选获取差异列:

row1 = cars[cars['Employee Name'] == 'John'].iloc[0]
row2 = cars[cars['Employee Name'] == 'Jill'].iloc[0]
diff_cols = row1[row1 != row2].index.tolist()
print(diff_cols)

输出结果:

['Sex', '# Cars Sold', 'Years With Company', 'Personal Car']

关键说明

  • 用iloc[0]可以直接提取筛选后的第一行数据(转为Series),避免返回带原索引的Series导致对比失败。
  • 上述方法支持所有数据类型的对比(字符串、数值、布尔等),解决了pd.diff仅支持数值型的局限。

内容的提问来源于stack exchange,提问作者E_Sarousi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 14:53:18