如何按章节标题与内容拆分超500页PDF并存入DataFrame用于NLP处理?
拆分PDF章节并提取内容到DataFrame的实现方案
针对500页PDF的章节拆分、标题内容提取需求,用Python工具链可以高效完成,以下是具体步骤和代码示例:
一、必备工具库
先安装核心依赖:
pip install pdfplumber pandas
pdfplumber:比PyPDF2更精准提取文本,支持字体、布局信息识别,适合处理复杂格式的PDFpandas:用来构建和管理title/content结构的DataFrame
如果是扫描版PDF(图片转的PDF),还需要安装OCR工具:
pip install pytesseract pillow
二、核心实现步骤
1. 提取PDF文本内容
遍历PDF每页,提取纯文本(扫描版则用OCR转换):
2. 识别章节标题
通过正则表达式匹配层级化标题(比如1.1.1 THE STANDARD FOR PROJECT MANAGEMENT这种格式),正则模式可以根据实际标题格式调整。
3. 拆分标题对应内容
遍历文本行,遇到新标题时保存上一个标题的内容,逐步收集所有章节的标题和对应内容。
4. 转换为DataFrame
把收集到的章节数据整理成pandas DataFrame,方便后续NLP处理。
三、完整代码示例
import pdfplumber import pandas as pd import re # 自定义正则匹配标题,可根据实际格式调整 # 匹配以数字层级(如1、1.1、1.1.1)开头的标题行 title_pattern = re.compile(r'^\d+(\.\d+)*\s+.+$', re.MULTILINE) def extract_pdf_chapters(pdf_file_path): chapter_list = [] current_title = None current_content_buffer = [] with pdfplumber.open(pdf_file_path) as pdf: for page in pdf.pages: # 提取当前页文本,扫描版PDF用page.to_image().to_string()替代 page_text = page.extract_text() if not page_text: continue # 按行拆分文本,逐行处理 for line in page_text.split('\n'): cleaned_line = line.strip() if not cleaned_line: continue # 匹配到新章节标题 if title_pattern.match(cleaned_line): # 先保存上一个章节的内容(如果存在) if current_title is not None: chapter_list.append({ 'title': current_title, 'content': '\n'.join(current_content_buffer).strip() }) # 更新当前标题,重置内容缓冲 current_title = cleaned_line current_content_buffer = [] else: # 非标题行,加入当前章节内容缓冲 if current_title is not None: current_content_buffer.append(cleaned_line) # 处理最后一个章节 if current_title is not None and current_content_buffer: chapter_list.append({ 'title': current_title, 'content': '\n'.join(current_content_buffer).strip() }) # 转换为DataFrame return pd.DataFrame(chapter_list) # 使用示例 if __name__ == '__main__': target_pdf = "your_500page_file.pdf" chapter_df = extract_pdf_chapters(target_pdf) # 查看前5条数据 print(chapter_df.head()) # 保存为CSV文件,方便后续NLP处理 chapter_df.to_csv("extracted_chapters.csv", index=False)
四、优化与适配建议
- 标题格式复杂:如果标题是加粗/大字号,可通过
page.extract_words()获取字体属性(如font_weight),筛选符合格式的行作为标题,避免正则误匹配。 - 大文件内存优化:500页PDF内存压力不大,若文件更大,可逐页处理后及时清理临时变量,或分批次写入CSV而非一次性存入列表。
- 正则调整:如果标题是
Chapter 1.1.1: XXX或带括号的格式,修改正则为r'^(Chapter\s)?\d+(\.\d+)*[:\s]+.+'即可。 - 扫描版PDF:确保已安装Tesseract OCR引擎(需单独下载安装,不是仅pip安装),然后将
page.extract_text()替换为page.to_image().to_string()。
内容的提问来源于stack exchange,提问作者Taher ZRIBI
相关产品推荐
相关产品推荐

