You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PDF章节标题并将页面标题存入数组?

提取PDF页面标题到数组的可行方案

针对你遇到的问题,核心是要基于文本的元数据(字号、位置、字体样式)而非仅首行来识别标题,以下是用Python的pdfplumber库实现的可靠方案,该库能精准提取文本的字体大小、位置等信息:

步骤1:安装依赖

pip install pdfplumber

步骤2:核心代码实现

import pdfplumber

def extract_pdf_titles(pdf_path):
    titles = []
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            # 提取当前页面所有带元数据的文本块
            text_chunks = page.extract_text_objects()
            if not text_chunks:
                titles.append(None)
                continue
            
            # 计算页面有效文本的平均字号(排除过小的页码类文本)
            valid_font_sizes = [chunk['size'] for chunk in text_chunks if chunk['size'] > 8]
            if not valid_font_sizes:
                titles.append(None)
                continue
            avg_font_size = sum(valid_font_sizes) / len(valid_font_sizes)
            
            # 筛选标题候选:字号显著大于平均水平,且位于页面顶部区域
            # 可根据你的PDF调整倍数和区域比例
            title_candidates = [
                chunk for chunk in text_chunks
                if chunk['size'] > avg_font_size * 1.5
                and chunk['top'] < page.height * 0.2
            ]
            
            # 确定最终标题
            if title_candidates:
                # 按字号降序,取最大的那个
                title_candidates.sort(key=lambda x: x['size'], reverse=True)
                titles.append(title_candidates[0]['text'].strip())
            else:
                # 无符合条件的候选时,取顶部区域字号最大的文本
                top_area_chunks = [chunk for chunk in text_chunks if chunk['top'] < page.height * 0.2]
                if top_area_chunks:
                    top_area_chunks.sort(key=lambda x: x['size'], reverse=True)
                    titles.append(top_area_chunks[0]['text'].strip())
                else:
                    titles.append(None)
    return titles

# 调用示例
document_titles = extract_pdf_titles("你的PDF文件路径.pdf")
print(document_titles)

优化调整建议

  • 阈值适配:根据你的PDF实际情况,调整avg_font_size * 1.5中的倍数(比如1.2或2),以及page.height * 0.2的顶部区域比例(比如0.15或0.25)
  • 加粗判断:如果标题是加粗样式,可在筛选条件中加入chunk['fontname'].lower().endswith('bold')(不同字体的加粗命名可能为Arial-Bold、TimesNewRoman-Bold等,需适配)
  • 干扰排除:可加入文本长度过滤(比如标题长度大于3),或关键词排除(比如过滤掉固定页眉的公司名、页码)
  • 多候选处理:如果页面有多个大字号文本,可优先选择最靠上的(按chunk['top']排序),或结合文本内容判断(比如是否包含“章节”“第X章”等关键词)

内容的提问来源于stack exchange,提问作者gofQ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 09:01:16