You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于标题与子标题提取PDF文本?无书签场景处理方案咨询

基于标题/子标题提取无书签PDF文本的方法

当PDF没有内置书签时,可通过文本格式特征、布局位置或内容结构来识别标题/子标题,进而提取对应内容,以下是具体实现方案:

方案1:利用字体格式差异识别标题

无书签的PDF中,标题通常具备明显格式特征(如字号更大、字体加粗),借助pdfplumber可提取文本的字体元数据,据此筛选标题并提取后续内容。

import pdfplumber

def extract_by_heading_format(pdf_path):
    extracted_data = []
    current_heading = None
    current_content = []
    
    # 根据目标PDF调整标题特征阈值
    HEADING_CRITERIA = {"min_size": 14, "is_bold": True}

    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            # 遍历每个字符的元数据
            for char in page.chars:
                text = char["text"]
                font_size = char["size"]
                is_bold = "Bold" in char["fontname"] or "bold" in char["fontname"]
                
                # 判断是否为标题
                if font_size >= HEADING_CRITERIA["min_size"] and is_bold:
                    # 保存上一个标题的内容
                    if current_heading:
                        extracted_data.append({
                            "heading": current_heading.strip(),
                            "content": "".join(current_content).strip()
                        })
                        current_content = []
                    current_heading = text
                else:
                    # 非标题内容追加到当前块
                    if current_heading:
                        current_content.append(text)
            # 处理每页最后一个标题的内容
            if current_heading:
                extracted_data.append({
                    "heading": current_heading.strip(),
                    "content": "".join(current_content).strip()
                })
                current_heading = None
                current_content = []
    return extracted_data

# 使用示例
result = extract_by_heading_format("target.pdf")
for item in result:
    print(f"标题: {item['heading']}")
    print(f"内容: {item['content']}\n")

提示:需先分析目标PDF的标题格式(可通过pdfplumber打印字符元数据查看),调整HEADING_CRITERIA中的字号、加粗规则。

方案2:结合页面布局定位标题

标题通常位于页面上部,可通过文本块的top坐标(PDF坐标原点在左下角,y值越大越靠上)辅助识别:

# 在方案1的标题判断逻辑中添加位置条件
page_height = page.height
if (font_size >= HEADING_CRITERIA["min_size"] 
    and is_bold 
    and char["top"] >= page_height * 0.7):  # 标题位于页面上70%区域
    # 执行标题识别逻辑

方案3:基于内容结构匹配标题

若标题无格式差异,但有固定结构(如1. 引言、2.1 实验设计),可通过正则表达式匹配:

import re

# 匹配层级化标题格式,可根据实际调整
heading_regex = re.compile(r"^\d+(\.\d+)*\s+.+$")

# 在遍历文本时,用正则判断
text_block = text.strip()
if heading_regex.match(text_block):
    # 执行标题识别逻辑

注意事项

  • 不同PDF格式差异大,需针对目标文件调整识别规则,没有通用的万能方案
  • 若为扫描版PDF(图片转PDF),需先通过OCR工具(如pytesseract)提取文本,再执行上述逻辑
  • pdfplumber的words方法可按单词提取元数据,适合不需要精细字符控制的场景

内容的提问来源于stack exchange,提问作者humera nikhat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 10:10:06