You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自动从PDF文件提取章节?批量提取年报管理层报告需求

从大量PDF年报提取管理层报告的实用方案

核心问题分析

你之前用正则匹配标题和后续章节只拿到目录,大概率是因为PDF里的目录标题和正文标题格式/位置重叠,正则没区分开正文区域和目录区域的内容。

具体解决思路

1. 先区分PDF的目录与正文区域

  • 先过滤页码范围:多数年报的目录集中在前几页,正文从固定页码(比如第5页)之后开始,可以先排除前N页内容,再在剩余正文里匹配章节。
  • 利用PDF书签定位:规范的年报PDF通常带书签(大纲),可以直接通过书签找到「Management Report」对应的页码区间,再提取该区间内容。用Python的pdfplumber实现的示例代码:
import pdfplumber

with pdfplumber.open("annual_report.pdf") as pdf:
    # 遍历书签定位目标章节
    for bookmark in pdf.bookmarks:
        if "Management Report" in bookmark["title"]:
            start_page = bookmark["page"]
            # 取下一个书签页码作为结束位置(无后续书签则取最后一页)
            next_idx = pdf.bookmarks.index(bookmark) + 1
            end_page = pdf.bookmarks[next_idx]["page"] - 1 if next_idx < len(pdf.bookmarks) else len(pdf.pages) - 1
            # 提取区间内的文本
            content = ""
            for page_num in range(start_page, end_page + 1):
                content += pdf.pages[page_num].extract_text()
            print(content)
            break

2. 优化正则匹配逻辑

如果PDF没有书签,调整正则规则避开目录:

  • 目录标题通常带后续页码(比如「Management Report 15」),正文标题无页码后缀,正则可匹配不带数字后缀的「Management Report」作为起始,到下一个一级标题(比如「Financial Statements」)结束。
  • 利用标题格式特征:正文章节标题一般字号更大、加粗,pdfplumber能提取文本的字体信息,先筛选出字号≥14且带Bold标识的文本,确认是正文标题后再开始提取。示例代码片段:
with pdfplumber.open("annual_report.pdf") as pdf:
    start_extract = False
    content = ""
    target_title = "Management Report"
    next_title = "Financial Statements"
    for page in pdf.pages:
        # 提取带格式的文本块
        for chunk in page.extract_words(extra_attrs=["fontname", "size"]):
            text = chunk["text"]
            # 匹配正文起始标题(加粗+大字号)
            if not start_extract and target_title in text and "Bold" in chunk["fontname"] and chunk["size"] >=14:
                start_extract = True
                content += text + "\n"
            # 匹配结束标题,停止提取
            elif start_extract and next_title in text and "Bold" in chunk["fontname"] and chunk["size"] >=14:
                start_extract = False
                break
            elif start_extract:
                content += text + " "
        if not start_extract:
            break
    print(content)

3. 批量处理效率优化

  • 用多线程/多进程加速:Python的concurrent.futures模块可以同时处理多个PDF,避免单线程逐个处理耗时过长。
  • 统一遍历规则:如果文件名有规律(比如「Company_2023_Annual_Report.pdf」),可以批量遍历目标文件夹,自动处理所有文件。

避坑提示

  • 扫描版PDF是图片格式,无法直接提取文本,需要先用OCR工具(比如pytesseract)转成可编辑文本,再按上述方法处理。
  • 不同公司年报标题可能有变体(比如「Management Discussion and Analysis」),可以把这些变体加入匹配列表,避免遗漏。

内容的提问来源于stack exchange,提问作者Limps

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 05:40:21