如何用Python提取PDF章节标题并将页面标题存入数组?
提取PDF页面标题到数组的可行方案
针对你遇到的问题,核心是要基于文本的元数据(字号、位置、字体样式)而非仅首行来识别标题,以下是用Python的pdfplumber库实现的可靠方案,该库能精准提取文本的字体大小、位置等信息:
步骤1:安装依赖
pip install pdfplumber
步骤2:核心代码实现
import pdfplumber def extract_pdf_titles(pdf_path): titles = [] with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: # 提取当前页面所有带元数据的文本块 text_chunks = page.extract_text_objects() if not text_chunks: titles.append(None) continue # 计算页面有效文本的平均字号(排除过小的页码类文本) valid_font_sizes = [chunk['size'] for chunk in text_chunks if chunk['size'] > 8] if not valid_font_sizes: titles.append(None) continue avg_font_size = sum(valid_font_sizes) / len(valid_font_sizes) # 筛选标题候选:字号显著大于平均水平,且位于页面顶部区域 # 可根据你的PDF调整倍数和区域比例 title_candidates = [ chunk for chunk in text_chunks if chunk['size'] > avg_font_size * 1.5 and chunk['top'] < page.height * 0.2 ] # 确定最终标题 if title_candidates: # 按字号降序,取最大的那个 title_candidates.sort(key=lambda x: x['size'], reverse=True) titles.append(title_candidates[0]['text'].strip()) else: # 无符合条件的候选时,取顶部区域字号最大的文本 top_area_chunks = [chunk for chunk in text_chunks if chunk['top'] < page.height * 0.2] if top_area_chunks: top_area_chunks.sort(key=lambda x: x['size'], reverse=True) titles.append(top_area_chunks[0]['text'].strip()) else: titles.append(None) return titles # 调用示例 document_titles = extract_pdf_titles("你的PDF文件路径.pdf") print(document_titles)
优化调整建议
- 阈值适配:根据你的PDF实际情况,调整
avg_font_size * 1.5中的倍数(比如1.2或2),以及page.height * 0.2的顶部区域比例(比如0.15或0.25) - 加粗判断:如果标题是加粗样式,可在筛选条件中加入
chunk['fontname'].lower().endswith('bold')(不同字体的加粗命名可能为Arial-Bold、TimesNewRoman-Bold等,需适配) - 干扰排除:可加入文本长度过滤(比如标题长度大于3),或关键词排除(比如过滤掉固定页眉的公司名、页码)
- 多候选处理:如果页面有多个大字号文本,可优先选择最靠上的(按
chunk['top']排序),或结合文本内容判断(比如是否包含“章节”“第X章”等关键词)
内容的提问来源于stack exchange,提问作者gofQ
相关产品推荐
相关产品推荐

