You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按章节标题与内容拆分超500页PDF并存入DataFrame用于NLP处理?

拆分PDF章节并提取内容到DataFrame的实现方案

针对500页PDF的章节拆分、标题内容提取需求,用Python工具链可以高效完成,以下是具体步骤和代码示例:

一、必备工具库

先安装核心依赖:

pip install pdfplumber pandas
  • pdfplumber:比PyPDF2更精准提取文本,支持字体、布局信息识别,适合处理复杂格式的PDF
  • pandas:用来构建和管理title/content结构的DataFrame

如果是扫描版PDF(图片转的PDF),还需要安装OCR工具:

pip install pytesseract pillow

二、核心实现步骤

1. 提取PDF文本内容

遍历PDF每页,提取纯文本(扫描版则用OCR转换):

2. 识别章节标题

通过正则表达式匹配层级化标题(比如1.1.1 THE STANDARD FOR PROJECT MANAGEMENT这种格式),正则模式可以根据实际标题格式调整。

3. 拆分标题对应内容

遍历文本行,遇到新标题时保存上一个标题的内容,逐步收集所有章节的标题和对应内容。

4. 转换为DataFrame

把收集到的章节数据整理成pandas DataFrame,方便后续NLP处理。

三、完整代码示例

import pdfplumber
import pandas as pd
import re

# 自定义正则匹配标题,可根据实际格式调整
# 匹配以数字层级(如1、1.1、1.1.1)开头的标题行
title_pattern = re.compile(r'^\d+(\.\d+)*\s+.+$', re.MULTILINE)

def extract_pdf_chapters(pdf_file_path):
    chapter_list = []
    current_title = None
    current_content_buffer = []

    with pdfplumber.open(pdf_file_path) as pdf:
        for page in pdf.pages:
            # 提取当前页文本,扫描版PDF用page.to_image().to_string()替代
            page_text = page.extract_text()
            if not page_text:
                continue
            
            # 按行拆分文本,逐行处理
            for line in page_text.split('\n'):
                cleaned_line = line.strip()
                if not cleaned_line:
                    continue
                
                # 匹配到新章节标题
                if title_pattern.match(cleaned_line):
                    # 先保存上一个章节的内容(如果存在)
                    if current_title is not None:
                        chapter_list.append({
                            'title': current_title,
                            'content': '\n'.join(current_content_buffer).strip()
                        })
                    # 更新当前标题,重置内容缓冲
                    current_title = cleaned_line
                    current_content_buffer = []
                else:
                    # 非标题行,加入当前章节内容缓冲
                    if current_title is not None:
                        current_content_buffer.append(cleaned_line)
        
        # 处理最后一个章节
        if current_title is not None and current_content_buffer:
            chapter_list.append({
                'title': current_title,
                'content': '\n'.join(current_content_buffer).strip()
            })
    
    # 转换为DataFrame
    return pd.DataFrame(chapter_list)

# 使用示例
if __name__ == '__main__':
    target_pdf = "your_500page_file.pdf"
    chapter_df = extract_pdf_chapters(target_pdf)
    
    # 查看前5条数据
    print(chapter_df.head())
    # 保存为CSV文件,方便后续NLP处理
    chapter_df.to_csv("extracted_chapters.csv", index=False)

四、优化与适配建议

  • 标题格式复杂:如果标题是加粗/大字号,可通过page.extract_words()获取字体属性(如font_weight),筛选符合格式的行作为标题,避免正则误匹配。
  • 大文件内存优化:500页PDF内存压力不大,若文件更大,可逐页处理后及时清理临时变量,或分批次写入CSV而非一次性存入列表。
  • 正则调整:如果标题是Chapter 1.1.1: XXX或带括号的格式,修改正则为r'^(Chapter\s)?\d+(\.\d+)*[:\s]+.+'即可。
  • 扫描版PDF:确保已安装Tesseract OCR引擎(需单独下载安装,不是仅pip安装),然后将page.extract_text()替换为page.to_image().to_string()。

内容的提问来源于stack exchange,提问作者Taher ZRIBI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 09:11:54