You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyMUPDF提取科研PDF文本,如何修正粘连以优化Spacy分句?

解决PDF科研文本提取后粘连内容的分句问题

问题场景

用PyMUPDF提取科研PDF文本并预处理后,不同逻辑区块(如ResearchGate提示、论文标题、会议信息)被空格粘连,导致Spacy无法正确分句。例如:

错误分句结果:
["See discussions, stats, and author profiles for this publication at: replaced_link. HDSKG:", "Harvesting Domain Specific Knowledge Graph from Content of Webpages Conference Paper · February 2017"]
期望分句结果:
["See discussions, stats, and author profiles for this publication at: replaced_link.", "HDSKG:Harvesting Domain Specific Knowledge Graph from Content of Webpages", "Conference Paper · February 2017"]

解决方案

通过针对性正则识别粘连边界,将空格替换为换行符,让Spacy能识别独立语义块。修改预处理代码,添加以下规则:

  • 处理URL替换后接大写缩写的情况
  • 拆分标题与会议/期刊信息的粘连
  • 识别独立大写缩写短语前的空格并替换为换行

修改后的完整代码

import fitz
import re

def extract_pdf_text(pdf_file_path):
    doc = fitz.open(pdf_file_path)
    text = ""
    for page in doc:
        text += page.get_text("text")
    return text

pdf_path = "/home/xxx/Papers/xxxxx.pdf"
text = extract_pdf_text(pdf_path)

# 基础预处理
text = re.sub(r"�", " ", text)
url_pattern = re.compile(r'http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\(\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+')
text = re.sub(url_pattern, 'replaced_link.', text)

# 新增粘连区块拆分规则
# 规则1:处理替换后的链接后接大写缩写的情况
text = re.sub(r'(replaced_link\.)\s+([A-Z]{2,}:)', r'\1\n\2', text)
# 规则2:处理标题与会议/期刊信息的粘连
text = re.sub(r'([a-zA-Z0-9])\s+(Conference Paper|Journal Article)', r'\1\n\2', text)
# 规则3:通用处理独立大写缩写短语前的空格
text = re.sub(r'\s+([A-Z]{2,}:)', r'\n\1', text)

# 最后合并多余空格
text = re.sub(r"\s+", " ", text)

效果说明

修改后,原粘连的文本会被拆分为带换行的独立区块:

"See discussions, stats, and author profiles for this publication at: replaced_link.\nHDSKG:Harvesting Domain Specific Knowledge Graph from Content of Webpages\nConference Paper · February 2017"

此时再用Spacy分句,就能得到符合预期的结果。

内容的提问来源于stack exchange,提问作者Faulheit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 07:15:26