You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从富格式PDF提取核心文本(排除页脚、表格等)

提取PDF核心文本的解决方案
  • 基于布局分析的pdfplumber进阶用法
    pdfplumber支持获取文本的位置、字体大小等布局特征,可通过这些特征过滤冗余内容:

    import pdfplumber
    
    def extract_main_text(pdf_path, page_num):
        main_text = []
        with pdfplumber.open(pdf_path) as pdf:
            page = pdf.pages[page_num]
            # 遍历页面所有字符块
            for word in page.extract_words():
                # 示例规则:过滤页脚(页面底部10%区域)、小字体文本
                if word['bottom'] > page.height * 0.9:
                    continue
                if word['size'] < 10:
                    continue
                main_text.append(word['text'])
        return ' '.join(main_text)
    
    # 使用示例
    print(extract_main_text("your_file.pdf", 10))
    

    可根据目标PDF的实际布局,调整坐标阈值、字体大小、字体样式等规则,精准筛选核心文本。

  • 基于机器学习的文档理解库

    • LayoutLM系列模型:针对文档布局的预训练模型,可识别文本的角色(正文、页脚、表格等),通过分类过滤非正文内容。
    • DocTR:专注文档理解的库,支持文本区域类型识别,直接提取正文:
      from doctr.io import DocumentFile
      from doctr.models import ocr_predictor
      
      model = ocr_predictor(pretrained=True)
      doc = DocumentFile.from_pdf("your_file.pdf")
      result = model(doc)
      # 提取模型标记为正文的内容(标签需根据模型输出调整)
      main_text = []
      for page in result.pages:
          for block in page.blocks:
              if block.label == "text":
                  main_text.append(block.value)
      print('\n'.join(main_text))
      
  • 商业工具API(可选)
    部分商业PDF处理API(如Adobe PDF Services API)内置文档结构识别功能,可直接提取正文,无需自行编写复杂规则,但需付费使用。

内容的提问来源于stack exchange,提问作者a-caputo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 18:25:37