You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas检测两个dataframe之间的差异及非法字符

方案结论

优先选用pandas实现,代码量更少,无需引入额外依赖,完全适配当前明确错误类型的校验场景。difflib更适合需要模糊匹配相似文本的需求,当前场景没必要额外引入。

最简实现代码

import pandas as pd

# 示例数据初始化
list_reference = ['hey', 'ola', 'bonjour', 'Hello']
list_test = ['Hey', 'ol a', "bon'jour", 'He"llo']
df_reference = pd.DataFrame(list_reference, columns=['content'])
df_test = pd.DataFrame(list_test, columns=['content'])

# 合并对齐待校验数据
df_compare = pd.DataFrame({
    '测试内容': df_test['content'],
    '参考内容': df_reference['content']
})

# 错误校验逻辑
def get_error_info(row):
    test_val = row['测试内容']
    ref_val = row['参考内容']
    errors = []
    if ' ' in test_val:
        errors.append('存在多余空格')
    if "'" in test_val:
        errors.append('存在单引号')
    if '"' in test_val:
        errors.append('存在双引号')
    if test_val.lower() == ref_val.lower() and test_val != ref_val:
        errors.append('大小写不合规')
    return '、'.join(errors) if errors else '匹配正常'

# 生成校验结果
df_compare['错误提示'] = df_compare.apply(get_error_info, axis=1)

校验结果示例

执行后输出df_compare即可看到每一行的错误标注:

测试内容参考内容错误提示
Heyhey大小写不合规
ol aola存在多余空格
bon'jourbonjour存在单引号
He"lloHello存在双引号

整个实现仅20行左右代码,基于pandas原生方法,批量处理大数据量时性能也优于difflib逐行比对的实现。

内容的提问来源于stack exchange,提问作者meuhfunk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 17:15:01