如何使用pandas检测两个dataframe之间的差异及非法字符
方案结论
优先选用pandas实现,代码量更少,无需引入额外依赖,完全适配当前明确错误类型的校验场景。difflib更适合需要模糊匹配相似文本的需求,当前场景没必要额外引入。
最简实现代码
import pandas as pd # 示例数据初始化 list_reference = ['hey', 'ola', 'bonjour', 'Hello'] list_test = ['Hey', 'ol a', "bon'jour", 'He"llo'] df_reference = pd.DataFrame(list_reference, columns=['content']) df_test = pd.DataFrame(list_test, columns=['content']) # 合并对齐待校验数据 df_compare = pd.DataFrame({ '测试内容': df_test['content'], '参考内容': df_reference['content'] }) # 错误校验逻辑 def get_error_info(row): test_val = row['测试内容'] ref_val = row['参考内容'] errors = [] if ' ' in test_val: errors.append('存在多余空格') if "'" in test_val: errors.append('存在单引号') if '"' in test_val: errors.append('存在双引号') if test_val.lower() == ref_val.lower() and test_val != ref_val: errors.append('大小写不合规') return '、'.join(errors) if errors else '匹配正常' # 生成校验结果 df_compare['错误提示'] = df_compare.apply(get_error_info, axis=1)
校验结果示例
执行后输出df_compare即可看到每一行的错误标注:
| 测试内容 | 参考内容 | 错误提示 |
|---|---|---|
| Hey | hey | 大小写不合规 |
| ol a | ola | 存在多余空格 |
| bon'jour | bonjour | 存在单引号 |
| He"llo | Hello | 存在双引号 |
整个实现仅20行左右代码,基于pandas原生方法,批量处理大数据量时性能也优于difflib逐行比对的实现。
内容的提问来源于stack exchange,提问作者meuhfunk
相关产品推荐
相关产品推荐

