如何用Python从富格式PDF提取核心文本(排除页脚、表格等)
提取PDF核心文本的解决方案
基于布局分析的pdfplumber进阶用法
pdfplumber支持获取文本的位置、字体大小等布局特征,可通过这些特征过滤冗余内容:import pdfplumber def extract_main_text(pdf_path, page_num): main_text = [] with pdfplumber.open(pdf_path) as pdf: page = pdf.pages[page_num] # 遍历页面所有字符块 for word in page.extract_words(): # 示例规则:过滤页脚(页面底部10%区域)、小字体文本 if word['bottom'] > page.height * 0.9: continue if word['size'] < 10: continue main_text.append(word['text']) return ' '.join(main_text) # 使用示例 print(extract_main_text("your_file.pdf", 10))可根据目标PDF的实际布局,调整坐标阈值、字体大小、字体样式等规则,精准筛选核心文本。
基于机器学习的文档理解库
- LayoutLM系列模型:针对文档布局的预训练模型,可识别文本的角色(正文、页脚、表格等),通过分类过滤非正文内容。
- DocTR:专注文档理解的库,支持文本区域类型识别,直接提取正文:
from doctr.io import DocumentFile from doctr.models import ocr_predictor model = ocr_predictor(pretrained=True) doc = DocumentFile.from_pdf("your_file.pdf") result = model(doc) # 提取模型标记为正文的内容(标签需根据模型输出调整) main_text = [] for page in result.pages: for block in page.blocks: if block.label == "text": main_text.append(block.value) print('\n'.join(main_text))
商业工具API(可选)
部分商业PDF处理API(如Adobe PDF Services API)内置文档结构识别功能,可直接提取正文,无需自行编写复杂规则,但需付费使用。
内容的提问来源于stack exchange,提问作者a-caputo
相关产品推荐
相关产品推荐

