如何使用Python提取PDF每页标题及高亮文本
问题原因
PyPDF2自带的extractText()方法仅能按读取顺序提取页面纯文本,无法获取字体字号、文本位置、标注等格式元数据,因此既无法区分标题与正文,也无法提取高亮标注内容,只能输出全页文本属于该库本身的能力局限。
无统一字号规则的单页标题提取方案
换用支持读取文本格式属性的PyMuPDF(导入名为fitz)实现,不需要预设全局标题字号,基于单页内的文本特征自动识别,核心识别规则:
- 标题默认出现在页面顶部1/4区域内,先排除页面中下部的正文内容
- 单页内字号越大、字体加粗的文本,为标题的概率越高(不依赖全局统一字号,仅在当前页做相对比较)
- 过滤无效候选:长度超过50字的长文本判定为正文段落,字号小于页面正文平均字号的短文本判定为页眉、页码
代码中的顶部区域占比、标题最大长度阈值可根据自身文档格式调整。
首先安装依赖:pip install pymupdf
实现代码:
import fitz def extract_page_titles(pdf_path): doc = fitz.open(pdf_path) page_titles = [] for page in doc: page_height = page.rect.height text_blocks = page.get_text("dict")["blocks"] candidate_blocks = [] # 统计当前页正文基准字号 font_sizes = [] for block in text_blocks: if "lines" not in block: continue for line in block["lines"]: for span in line["spans"]: if span["text"].strip(): font_sizes.append(span["size"]) base_font_size = sum(font_sizes)/len(font_sizes) if font_sizes else 10 # 筛选顶部区域的候选标题 for block in text_blocks: if "lines" not in block or block["bbox"][1] > page_height/4: continue block_text = "" max_size = 0 is_bold = False for line in block["lines"]: for span in line["spans"]: block_text += span["text"] if span["size"] > max_size: max_size = span["size"] if "bold" in span["font"].lower(): is_bold = True block_text = block_text.strip() # 过滤不符合规则的候选 if not block_text or len(block_text) > 50 or max_size < base_font_size: continue # 计算匹配权重 weight = max_size * (2 if is_bold else 1) candidate_blocks.append((weight, block_text)) # 取权重最高的作为当前页标题 if candidate_blocks: candidate_blocks.sort(reverse=True, key=lambda x:x[0]) page_titles.append(candidate_blocks[0][1]) else: page_titles.append("未识别到标题") doc.close() return page_titles # 调用示例 titles = extract_page_titles("Test2.pdf") for idx, title in enumerate(titles, 1): print(f"第{idx}页标题:{title}")
PDF高亮文本提取方法
同样基于PyMuPDF实现,高亮属于PDF的内置注释类型,直接遍历页面注释筛选高亮类注释,提取关联区域的文本即可,实现代码:
import fitz def extract_highlight_text(pdf_path): doc = fitz.open(pdf_path) all_highlights = [] for page_num, page in enumerate(doc, 1): page_highlights = [] annot = page.first_annot while annot: # 8对应PyMuPDF中定义的高亮注释类型 if annot.type[0] == 8: highlight_text = page.get_textbox(annot.rect).strip() if highlight_text: page_highlights.append(highlight_text) annot = annot.next all_highlights.append({ "page": page_num, "highlights": page_highlights }) doc.close() return all_highlights # 调用示例 highlight_result = extract_highlight_text("Test2.pdf") for item in highlight_result: if item["highlights"]: print(f"第{item['page']}页高亮内容:") for h in item["highlights"]: print(f"- {h}")
注意:如果待处理PDF是扫描生成的图片版PDF,以上方法无法直接识别文本和标题,需要先做OCR处理后再匹配规则。
内容的提问来源于stack exchange,提问作者Prajkta Mangulkar
相关产品推荐
相关产品推荐

