如何基于标题与子标题提取PDF文本?无书签场景处理方案咨询
基于标题/子标题提取无书签PDF文本的方法
当PDF没有内置书签时,可通过文本格式特征、布局位置或内容结构来识别标题/子标题,进而提取对应内容,以下是具体实现方案:
方案1:利用字体格式差异识别标题
无书签的PDF中,标题通常具备明显格式特征(如字号更大、字体加粗),借助pdfplumber可提取文本的字体元数据,据此筛选标题并提取后续内容。
import pdfplumber def extract_by_heading_format(pdf_path): extracted_data = [] current_heading = None current_content = [] # 根据目标PDF调整标题特征阈值 HEADING_CRITERIA = {"min_size": 14, "is_bold": True} with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: # 遍历每个字符的元数据 for char in page.chars: text = char["text"] font_size = char["size"] is_bold = "Bold" in char["fontname"] or "bold" in char["fontname"] # 判断是否为标题 if font_size >= HEADING_CRITERIA["min_size"] and is_bold: # 保存上一个标题的内容 if current_heading: extracted_data.append({ "heading": current_heading.strip(), "content": "".join(current_content).strip() }) current_content = [] current_heading = text else: # 非标题内容追加到当前块 if current_heading: current_content.append(text) # 处理每页最后一个标题的内容 if current_heading: extracted_data.append({ "heading": current_heading.strip(), "content": "".join(current_content).strip() }) current_heading = None current_content = [] return extracted_data # 使用示例 result = extract_by_heading_format("target.pdf") for item in result: print(f"标题: {item['heading']}") print(f"内容: {item['content']}\n")
提示:需先分析目标PDF的标题格式(可通过
pdfplumber打印字符元数据查看),调整HEADING_CRITERIA中的字号、加粗规则。
方案2:结合页面布局定位标题
标题通常位于页面上部,可通过文本块的top坐标(PDF坐标原点在左下角,y值越大越靠上)辅助识别:
# 在方案1的标题判断逻辑中添加位置条件 page_height = page.height if (font_size >= HEADING_CRITERIA["min_size"] and is_bold and char["top"] >= page_height * 0.7): # 标题位于页面上70%区域 # 执行标题识别逻辑
方案3:基于内容结构匹配标题
若标题无格式差异,但有固定结构(如1. 引言、2.1 实验设计),可通过正则表达式匹配:
import re # 匹配层级化标题格式,可根据实际调整 heading_regex = re.compile(r"^\d+(\.\d+)*\s+.+$") # 在遍历文本时,用正则判断 text_block = text.strip() if heading_regex.match(text_block): # 执行标题识别逻辑
注意事项
- 不同PDF格式差异大,需针对目标文件调整识别规则,没有通用的万能方案
- 若为扫描版PDF(图片转PDF),需先通过OCR工具(如pytesseract)提取文本,再执行上述逻辑
pdfplumber的words方法可按单词提取元数据,适合不需要精细字符控制的场景
内容的提问来源于stack exchange,提问作者humera nikhat
相关产品推荐
相关产品推荐

