忽略文本前编号,比较Pandas DataFrame列内容
忽略编号仅对比文本内容的DataFrame列差异检测
你当前的代码直接对比完整字符串,导致编号不同但后续文本一致的行被误判为差异。要实现忽略编号仅对比文本部分,核心是先提取编号后的纯文本,再进行比较。
修改后的代码方案
import pandas as pd import numpy as np if __name__ == "__main__": data1 = pd.DataFrame({ "ID": [1, 2, 3, 4], "Text": ['1-1 Text here1', '1-2 Text here2', '1-3 Text here3', '1-4 Text here4']}) data2 = pd.DataFrame({ "ID": [1, 2, 3, 4], "Text": ['1-1 Text here1', '1-4 Text here', '1-5 Text here3', '1-6 Text here']}) df_both = pd.concat([data1.set_index('ID'), data2.set_index('ID')], axis=1, keys=['Data1', 'Data2']) df_both = df_both.swaplevel(axis='columns') # 提取编号后的纯文本:分割第一个空格后的内容 def get_clean_text(text_str): return text_str.split(' ', 1)[1] if ' ' in text_str else text_str for column in df_both.columns.levels[0]: # 分别提取两数据源的纯文本 clean_data1 = df_both[column]['Data1'].apply(get_clean_text) clean_data2 = df_both[column]['Data2'].apply(get_clean_text) # 对比纯文本,找出差异ID diff_indices = np.where(clean_data1 != clean_data2) print(df_both.index[diff_indices])
运行结果
Int64Index([2, 4], dtype='int64', name='ID')
关键说明
- 函数
get_clean_text通过split(' ', 1)将字符串按第一个空格拆分,直接获取编号后的文本内容,快速剥离X-X格式的编号前缀。 - 如果编号格式有变化(比如多空格、特殊符号),可以改用正则表达式实现更精准的匹配:
这个正则会严格匹配import re def get_clean_text(text_str): match_result = re.match(r'^\d+-\d+\s+(.*)', text_str) return match_result.group(1) if match_result else text_str数字-数字 + 空格开头的前缀,确保只提取后续的文本部分。
内容的提问来源于stack exchange,提问作者Anna Ortega
相关产品推荐
相关产品推荐

