You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PymuPDF解析PDF时内容对齐与逻辑分组失败求助

PyMuPDF解析PDF章节内容错乱的解决方法

问题根源

你当前的代码仅通过文本标点和硬编码关键词来分组内容,没有利用PDF的布局结构信息(比如标题与内容的位置、字号差异),导致排版换行被误判成语义换行,进而出现文本错乱拼接的问题。硬编码主题的方式也无法适应PDF排版的变化,很容易匹配失败。

改进思路

  1. 利用块的布局属性区分标题与内容:PDF中章节标题通常字号更大、位置更靠上,可通过文本块的size(字号)和bbox(位置)来识别。
  2. 按页面顺序排序文本块:PyMuPDF返回的块不一定是从上到下的顺序,需按块的顶部y坐标排序,保证处理顺序与阅读顺序一致。
  3. 基于标题关联对应内容:识别到标题后,将后续的内容块归到该标题下,直到遇到下一个标题为止。

完整实现代码

import requests
import pymupdf

url = "https://www.iipa.org.in/upload/IPG_const.pdf"
response = requests.get(url)
doc = pymupdf.open(stream=response.content, filetype="pdf")

page = doc[24]  # 对应第25页
blocks = page.get_text("dict")["blocks"]

# 过滤出文本块,排除图片等非文本内容
text_blocks = []
for block in blocks:
    if block["type"] == 0:  # type=0表示文本块
        # 提取块的基本信息:文本、字号、顶部y坐标
        lines = block.get("lines", [])
        if lines:
            # 取第一行的字号作为块的字号(标题通常整行字号一致)
            base_font_size = lines[0]["spans"][0]["size"]
            # 块的顶部y坐标,用于排序
            top_y = block["bbox"][1]
            # 拼接块内所有文本
            block_text = " ".join([span["text"] for line in lines for span in line["spans"]]).strip()
            text_blocks.append({
                "text": block_text,
                "font_size": base_font_size,
                "top_y": top_y
            })

# 按顶部y坐标升序排序,保证从上到下处理文本块
text_blocks.sort(key=lambda x: x["top_y"])

# 定义标题字号阈值(根据实际PDF调整,这里取12作为分界)
TITLE_FONT_THRESHOLD = 12
sections = []
current_section = {"title": None, "content": ""}

for block in text_blocks:
    text = block["text"]
    font_size = block["font_size"]
    
    # 判断是否为标题:字号大于阈值,且文本是目标主题(也可仅靠字号判断)
    is_title = font_size > TITLE_FONT_THRESHOLD or text.lower() in ["liberty", "justice", "equality", "fraternity", "social economic political", "distributive justice"]
    
    if is_title:
        # 保存上一个章节(如果有内容)
        if current_section["title"] is not None and current_section["content"].strip():
            sections.append(current_section)
        # 开始新章节
        current_section = {"title": text, "content": ""}
    else:
        # 拼接内容,处理排版换行(直接追加,因为PDF换行是排版需要)
        if current_section["content"]:
            current_section["content"] += " " + text
        else:
            current_section["content"] = text

# 保存最后一个章节
if current_section["title"] is not None and current_section["content"].strip():
    sections.append(current_section)

# 输出结果
for section in sections:
    print(f"*标题:* {section['title']}")
    print(f"*内容:* {section['content']}")
    print("-----")

关键优化点

  • 布局识别:通过字号和位置判断标题,比硬编码关键词更鲁棒,即使PDF排版微调也能适配。
  • 顺序修正:对文本块按顶部坐标排序,确保处理顺序与页面阅读顺序一致。
  • 内容拼接:直接追加内容块文本,避免因排版换行导致的语义断裂,解决“社会正义”部分的文本错乱问题。

内容的提问来源于stack exchange,提问作者Abinash abi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 15:33:16