使用PymuPDF解析PDF时内容对齐与逻辑分组失败求助
PyMuPDF解析PDF章节内容错乱的解决方法
问题根源
你当前的代码仅通过文本标点和硬编码关键词来分组内容,没有利用PDF的布局结构信息(比如标题与内容的位置、字号差异),导致排版换行被误判成语义换行,进而出现文本错乱拼接的问题。硬编码主题的方式也无法适应PDF排版的变化,很容易匹配失败。
改进思路
- 利用块的布局属性区分标题与内容:PDF中章节标题通常字号更大、位置更靠上,可通过文本块的
size(字号)和bbox(位置)来识别。 - 按页面顺序排序文本块:PyMuPDF返回的块不一定是从上到下的顺序,需按块的顶部y坐标排序,保证处理顺序与阅读顺序一致。
- 基于标题关联对应内容:识别到标题后,将后续的内容块归到该标题下,直到遇到下一个标题为止。
完整实现代码
import requests import pymupdf url = "https://www.iipa.org.in/upload/IPG_const.pdf" response = requests.get(url) doc = pymupdf.open(stream=response.content, filetype="pdf") page = doc[24] # 对应第25页 blocks = page.get_text("dict")["blocks"] # 过滤出文本块,排除图片等非文本内容 text_blocks = [] for block in blocks: if block["type"] == 0: # type=0表示文本块 # 提取块的基本信息:文本、字号、顶部y坐标 lines = block.get("lines", []) if lines: # 取第一行的字号作为块的字号(标题通常整行字号一致) base_font_size = lines[0]["spans"][0]["size"] # 块的顶部y坐标,用于排序 top_y = block["bbox"][1] # 拼接块内所有文本 block_text = " ".join([span["text"] for line in lines for span in line["spans"]]).strip() text_blocks.append({ "text": block_text, "font_size": base_font_size, "top_y": top_y }) # 按顶部y坐标升序排序,保证从上到下处理文本块 text_blocks.sort(key=lambda x: x["top_y"]) # 定义标题字号阈值(根据实际PDF调整,这里取12作为分界) TITLE_FONT_THRESHOLD = 12 sections = [] current_section = {"title": None, "content": ""} for block in text_blocks: text = block["text"] font_size = block["font_size"] # 判断是否为标题:字号大于阈值,且文本是目标主题(也可仅靠字号判断) is_title = font_size > TITLE_FONT_THRESHOLD or text.lower() in ["liberty", "justice", "equality", "fraternity", "social economic political", "distributive justice"] if is_title: # 保存上一个章节(如果有内容) if current_section["title"] is not None and current_section["content"].strip(): sections.append(current_section) # 开始新章节 current_section = {"title": text, "content": ""} else: # 拼接内容,处理排版换行(直接追加,因为PDF换行是排版需要) if current_section["content"]: current_section["content"] += " " + text else: current_section["content"] = text # 保存最后一个章节 if current_section["title"] is not None and current_section["content"].strip(): sections.append(current_section) # 输出结果 for section in sections: print(f"*标题:* {section['title']}") print(f"*内容:* {section['content']}") print("-----")
关键优化点
- 布局识别:通过字号和位置判断标题,比硬编码关键词更鲁棒,即使PDF排版微调也能适配。
- 顺序修正:对文本块按顶部坐标排序,确保处理顺序与页面阅读顺序一致。
- 内容拼接:直接追加内容块文本,避免因排版换行导致的语义断裂,解决“社会正义”部分的文本错乱问题。
内容的提问来源于stack exchange,提问作者Abinash abi
相关产品推荐
相关产品推荐

