You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PDF特定章节提取异常求助:标准化章节搜索方案

标准化PDF第2章节(健康分类)搜索的解决方案

针对不同PDF格式下章节标题变体导致的匹配失败问题,可通过以下几种实用方法标准化搜索逻辑:

1. 构建正则表达式覆盖所有标题变体

把所有可能的章节标题变体整理成正则规则,覆盖中英文、数字/中文章节号、标题后缀等情况:

import re

def extract_healthclassifications(text):
    # 匹配第2章标题的所有常见变体
    chapter_pattern = re.compile(
        r'^(?:第?二|2)(?:章|\.| |\t|HEALTH)?\s*(?:健康分类|HEALTH CLASSIFICATIONS|健康分类.*)',
        re.IGNORECASE | re.MULTILINE
    )
    # 匹配下一章标题(用于确定当前章节的结束位置)
    next_chapter_pattern = re.compile(
        r'^(?:第?三|3)(?:章|\.| |\t|HEALTH)?\s*(?:.*)',
        re.IGNORECASE | re.MULTILINE
    )
    
    # 找到第2章的起始位置
    start_match = chapter_pattern.search(text)
    if not start_match:
        return ""
    
    start_pos = start_match.end()
    # 找到下一章的起始位置作为结束点
    end_match = next_chapter_pattern.search(text, start_pos)
    end_pos = end_match.start() if end_match else len(text)
    
    return text[start_pos:end_pos].strip()

2. 利用PDF结构化大纲(书签)定位章节

如果PDF包含大纲/书签,直接遍历大纲获取第2章的页码范围,比纯文本匹配更可靠:

import fitz  # PyMuPDF

def extract_healthclassifications_from_outline(pdf_path):
    doc = fitz.open(pdf_path)
    target_chapter = None
    # 遍历大纲找第2章
    for idx, item in enumerate(doc.outline):
        title = item.get('title', '').strip()
        if re.match(r'^(?:第?二|2)', title, re.IGNORECASE):
            target_chapter = item
            break
    if not target_chapter:
        return ""
    
    # 确定章节的页码范围
    start_page_num = doc[target_chapter.page].number
    # 找下一个章节的起始页
    end_page_num = len(doc)
    if idx + 1 < len(doc.outline):
        next_item = doc.outline[idx + 1]
        end_page_num = doc[next_item.page].number
    
    # 提取章节文本
    chapter_text = ""
    for page_num in range(start_page_num, end_page_num):
        chapter_text += doc[page_num].get_text()
    return chapter_text.strip()

3. 结合文本格式特征识别章节标题

章节标题通常有字号大、加粗、单独成段的格式特征,可通过提取文本块的排版信息辅助匹配:

import fitz

def extract_healthclassifications_by_format(pdf_path):
    doc = fitz.open(pdf_path)
    chapter_text = ""
    in_target_chapter = False
    target_keywords = ["健康分类", "HEALTH CLASSIFICATIONS"]
    
    for page in doc:
        blocks = page.get_text("dict")["blocks"]
        for block in blocks:
            if "lines" not in block:
                continue
            for line in block["lines"]:
                line_text = ''.join([span["text"] for span in line["spans"]]).strip()
                if not line_text:
                    continue
                
                # 检查是否是章节标题:字号≥14,加粗(flags&2表示加粗)
                span = line["spans"][0]
                is_title = span["size"] >= 14 and (span["flags"] & 2)
                # 匹配第2章的标题变体
                if is_title and re.match(r'^(?:第?二|2)', line_text, re.IGNORECASE):
                    # 检查是否包含目标关键字
                    if any(keyword in line_text.upper() for keyword in [k.upper() for k in target_keywords]):
                        in_target_chapter = True
                        continue
                
                # 进入目标章节后,遇到下一章标题则停止
                if in_target_chapter and is_title and re.match(r'^(?:第?三|3)', line_text, re.IGNORECASE):
                    in_target_chapter = False
                    break
                
                if in_target_chapter:
                    chapter_text += line_text + "\n"
    return chapter_text.strip()

4. 模糊匹配兜底(应对非标准化标题)

对于完全无规则的标题,可使用模糊匹配工具允许一定相似度的匹配:

from fuzzywuzzy import fuzz

def is_target_chapter_title(title, target_keywords, threshold=80):
    """检查标题是否匹配目标章节"""
    title_upper = title.upper()
    for keyword in target_keywords:
        if fuzz.partial_ratio(title_upper, keyword.upper()) >= threshold:
            return True
    return False

整合建议

实际使用时可按优先级组合方法:

  • 优先尝试大纲定位法,准确率最高;
  • 无大纲时,使用格式+正则匹配法;
  • 前两种都失败时,用模糊匹配+正则兜底。

内容的提问来源于stack exchange,提问作者Nuno Braga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 09:00:04