You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas报错TypeError: unhashable type: 'Series',如何筛选Return行找重复值?

仅标记Return行内重复Application number的解决方案

问题说明

需要在DataFrame中仅针对Document type为Return的行,找出重复的Application number,并在对应行的Comment列添加注释//Duplicate is found。现有代码运行时触发TypeError: unhashable type: 'Series'错误,若全局查找重复值则会错误标记Sale行的重复项。

原DataFrame示例

Document type   Application number
0       Return    1658
1       Sale      1658
2       Return    1659
3       Sale      1659
4       Return    1659
5       Return    1660
6       Return    1660

期望结果

Document type   Application number    Comment
0       Return    1658                  
1       Sale      1658
2       Return    1659                  //Duplicate is found
3       Sale      1659
4       Return    1659                  //Duplicate is found
5       Return    1660                  //Duplicate is found
6       Return    1660                  //Duplicate is found

报错代码

def check_duplicated_app_nums(df,
                              col_app_num,
                              col_doc_type,
                              col_comments,
                              comment = 'Duplicate is found'):

    mask_doc_type = df[col_doc_type] == 'Return'
    mask_duplicate = df[mask_doc_type].duplicated(subset=col_app_num, keep=False)

    df.loc[mask_duplicate, col_comments] = df.apply(lambda x: '%s//%s' % (x[col_comments], comment), axis=1)

错误原因

df[mask_doc_type].duplicated(...)返回的是仅包含Return行的Series,其索引仅对应原DataFrame中的部分行。直接将这个Series传入df.loc时,虽然索引能匹配,但后续的df.apply是对整个DataFrame执行操作,且掩码长度与原DataFrame不一致,最终触发类型错误。同时这种写法效率较低,没必要对全表执行apply。

修正方案

方法一:先定位重复的Application number,再批量标记

先从Return行中找出所有重复的Application number,再通过掩码匹配原DataFrame中符合条件的行,批量添加注释:

import pandas as pd

def check_duplicated_app_nums(df,
                              col_app_num,
                              col_doc_type,
                              col_comments,
                              comment='Duplicate is found'):
    # 复制原DataFrame避免修改原始数据(可选,按需调整)
    df = df.copy()
    # 筛选所有Return行
    return_subset = df[df[col_doc_type] == 'Return']
    # 找出Return行中重复的Application number集合
    duplicate_apps = return_subset[return_subset.duplicated(subset=col_app_num, keep=False)][col_app_num].unique()
    # 构建全表掩码:是Return行 且 Application number在重复集合中
    target_mask = (df[col_doc_type] == 'Return') & (df[col_app_num].isin(duplicate_apps))
    # 给目标行添加注释,兼容原有Comment不为空的情况
    df.loc[target_mask, col_comments] = df.loc[target_mask, col_comments].apply(
        lambda x: f"{x}//{comment}" if pd.notna(x) else f"//{comment}"
    )
    return df

方法二:对齐重复掩码到原DataFrame索引

先在Return行子集内生成重复标记,再将掩码扩展到原DataFrame的完整索引后赋值:

import pandas as pd

def check_duplicated_app_nums(df,
                              col_app_num,
                              col_doc_type,
                              col_comments,
                              comment='Duplicate is found'):
    df = df.copy()
    # 标记所有Return行
    mask_return = df[col_doc_type] == 'Return'
    # 在Return行内找出重复项的掩码
    duplicate_in_return = df[mask_return].duplicated(subset=col_app_num, keep=False)
    # 将掩码对齐到原DataFrame的完整索引
    full_duplicate_mask = pd.Series(False, index=df.index)
    full_duplicate_mask.loc[duplicate_in_return.index] = duplicate_in_return
    # 赋值注释
    df.loc[full_duplicate_mask, col_comments] = df.loc[full_duplicate_mask, col_comments].apply(
        lambda x: f"{x}//{comment}" if pd.notna(x) else f"//{comment}"
    )
    return df

测试验证

用示例数据测试函数:

# 构造示例数据
sample_data = {
    'Document type': ['Return', 'Sale', 'Return', 'Sale', 'Return', 'Return', 'Return'],
    'Application number': [1658, 1658, 1659, 1659, 1659, 1660, 1660]
}
df_test = pd.DataFrame(sample_data)
df_test['Comment'] = ''  # 初始化Comment列

# 调用函数
result = check_duplicated_app_nums(df_test, 'Application number', 'Document type', 'Comment')
print(result)

输出结果与期望一致。

内容的提问来源于stack exchange,提问作者Angelina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 10:13:16