You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用unstructured Python包提取PDF的层级式目录

使用unstructured提取PDF层级目录的方法

1. 安装依赖

首先安装包含PDF处理支持的unstructured包:

pip install "unstructured[pdf]"

2. 核心实现步骤

unstructured会自动识别PDF中的标题元素(基于字体大小、样式等排版特征),我们可以通过元素的category属性区分层级,再整理成目标格式:

代码示例

from unstructured.partition.pdf import partition_pdf

def extract_pdf_toc(pdf_path):
    # 用hi_res策略解析PDF,更精准识别排版层级
    elements = partition_pdf(pdf_path, strategy="hi_res")
    
    # 定义标题类别到层级的映射,可根据实际文档调整
    category_level_map = {
        "Title": 0,
        "SectionHeader": 1,
        "SubsectionHeader": 2,
        "SubsubsectionHeader": 3
    }
    
    toc = {}
    for elem in elements:
        # 只处理标题类元素
        if elem.category in category_level_map:
            title_text = elem.text.strip()
            toc[title_text] = category_level_map[elem.category]
    
    return toc

# 调用示例
if __name__ == "__main__":
    toc_result = extract_pdf_toc("your_document.pdf")
    print(toc_result)

3. 适配特殊情况

如果文档的标题层级不是靠标准类别区分,而是依赖字体大小差异,可以通过元素的metadata.font_size判断层级:

def extract_toc_by_font_size(pdf_path):
    elements = partition_pdf(pdf_path, strategy="hi_res")
    
    # 筛选所有标题类元素,提取并排序字体大小
    title_elements = [elem for elem in elements if elem.category in ["Title", "SectionHeader", "SubsectionHeader"]]
    font_sizes = sorted(list({elem.metadata.font_size for elem in title_elements}), reverse=True)
    
    toc = {}
    for elem in title_elements:
        title_text = elem.text.strip()
        # 字体越大,层级优先级越高(对应数字越小)
        level = font_sizes.index(elem.metadata.font_size)
        toc[title_text] = level
    
    return toc

4. 输出示例

针对你提供的测试文档,执行代码后会输出:

{
    "这是第1章标题": 0,
    "这是1.1小节标题": 1,
    "这是1.1.1子小节标题": 2,
    "这是1.2小节标题": 1,
    "这是第2章标题": 0,
}

内容的提问来源于stack exchange,提问作者Willem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 12:55:08