Python in运算符无法识别文本中已存在子串的问题排查
问题原因
匹配失败的核心是仅处理了普通半角空格,没有清理PDF提取文本中残留的换行及其他不可见空白字符。
你从PDF提取的目标文本中,对应标题被硬换行拆成了两行:
Synthesis, properties and applications of Janus nanoparticles
执行replace(" ", "")只会删除普通空格,两行之间的换行符\n会被完整保留,处理后的目标字符串对应片段为synthesis,propertiesandapplicationsofjanus\nnanoparticles,而处理后的待匹配子串为synthesis,propertiesandapplicationsofjanusnanoparticles,二者中间差了一个换行符,in运算符自然会返回False。
除此之外,用fitz(PyMuPDF)提取PDF文本时,还经常会夹带回车符\r、不间断空格\xa0、零宽空格等肉眼无法识别的特殊空白字符,仅替换普通空格无法覆盖所有场景。
修复方法
不要单独替换普通空格,改用正则匹配所有类型的空白字符统一移除,再做归一化匹配,修正后的代码如下:
from unidecode import unidecode import re example_string = '''available at www.sciencedirect.com journal homepage: www.elsevier.com/locate/nanotoday REVIEW Synthesis, properties and applications of Janus nanoparticles Marco Lattuada a, T. Alan Hatton b,''' list_of_titles = ["Synthesis, properties and applications of Janus nanoparticles", "another_title", "another_title"] def normalize_text(text): text = text.casefold() # 匹配所有空白类字符统一删除,覆盖换行、回车、制表符、不间断空格等特殊空白 text = re.sub(r'\s+', '', text) text = unidecode(text) return text normalized_target = normalize_text(example_string) for title in list_of_titles: if normalize_text(title) in normalized_target: print("Yes")
运行代码后会正常输出Yes。
补充提示
正则规则中的\s可以匹配所有Unicode定义的空白类字符,比手动逐个替换特殊字符的兼容性强很多。如果后续仍出现匹配失败的情况,可以将两边归一化后的字符串打印出来逐字符比对,针对性过滤PDF提取带出的特殊控制字符即可。
内容的提问来源于stack exchange,提问作者double_wizz
相关产品推荐
相关产品推荐

