PyMUPDF提取科研PDF文本,如何修正粘连以优化Spacy分句?
问题场景
用PyMUPDF提取科研PDF文本并预处理后,不同逻辑区块(如ResearchGate提示、论文标题、会议信息)被空格粘连,导致Spacy无法正确分句。例如:
错误分句结果:
["See discussions, stats, and author profiles for this publication at: replaced_link. HDSKG:", "Harvesting Domain Specific Knowledge Graph from Content of Webpages Conference Paper · February 2017"]
期望分句结果:
["See discussions, stats, and author profiles for this publication at: replaced_link.", "HDSKG:Harvesting Domain Specific Knowledge Graph from Content of Webpages", "Conference Paper · February 2017"]
解决方案
通过针对性正则识别粘连边界,将空格替换为换行符,让Spacy能识别独立语义块。修改预处理代码,添加以下规则:
- 处理URL替换后接大写缩写的情况
- 拆分标题与会议/期刊信息的粘连
- 识别独立大写缩写短语前的空格并替换为换行
修改后的完整代码
import fitz import re def extract_pdf_text(pdf_file_path): doc = fitz.open(pdf_file_path) text = "" for page in doc: text += page.get_text("text") return text pdf_path = "/home/xxx/Papers/xxxxx.pdf" text = extract_pdf_text(pdf_path) # 基础预处理 text = re.sub(r"�", " ", text) url_pattern = re.compile(r'http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\(\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+') text = re.sub(url_pattern, 'replaced_link.', text) # 新增粘连区块拆分规则 # 规则1:处理替换后的链接后接大写缩写的情况 text = re.sub(r'(replaced_link\.)\s+([A-Z]{2,}:)', r'\1\n\2', text) # 规则2:处理标题与会议/期刊信息的粘连 text = re.sub(r'([a-zA-Z0-9])\s+(Conference Paper|Journal Article)', r'\1\n\2', text) # 规则3:通用处理独立大写缩写短语前的空格 text = re.sub(r'\s+([A-Z]{2,}:)', r'\n\1', text) # 最后合并多余空格 text = re.sub(r"\s+", " ", text)
效果说明
修改后,原粘连的文本会被拆分为带换行的独立区块:
"See discussions, stats, and author profiles for this publication at: replaced_link.\nHDSKG:Harvesting Domain Specific Knowledge Graph from Content of Webpages\nConference Paper · February 2017"
此时再用Spacy分句,就能得到符合预期的结果。
内容的提问来源于stack exchange,提问作者Faulheit

