You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PDF法律合同的章节拆分至独立文本文件?

拆分PDF法律合同章节为独立文本文件的可行方案

方法一:Python脚本直接提取并拆分(推荐)

适合结构化清晰的PDF合同,通过正则匹配章节标题实现精准拆分:

  1. 安装依赖
pip install PyPDF2
  1. 运行以下脚本(根据合同标题格式调整正则)
import PyPDF2
import re

def split_contract_chapters(pdf_path):
    # 读取PDF全量文本
    with open(pdf_path, 'rb') as file:
        reader = PyPDF2.PdfReader(file)
        full_text = ""
        for page in reader.pages:
            full_text += page.extract_text() + "\n"
    
    # 匹配章节标题(示例:匹配“第X章 标题”格式,可按需修改正则)
    chapter_pattern = re.compile(r'(第\d+章\s+.+?)(?=第\d+章|$)', re.DOTALL)
    chapters = chapter_pattern.findall(full_text)
    
    # 生成独立文本文件
    for idx, chapter in enumerate(chapters, 1):
        chapter_clean = chapter.strip()
        # 提取标题作为文件名
        title_match = re.match(r'第\d+章\s+(.+)', chapter_clean.split('\n')[0])
        filename = f"章节{idx}_{title_match.group(1)}.txt" if title_match else f"章节{idx}.txt"
        with open(filename, 'w', encoding='utf-8') as f:
            f.write(chapter_clean)
        print(f"已生成:{filename}")

# 替换为你的PDF路径
split_contract_chapters("法律合同.pdf")

方法二:转HTML后解析拆分(你之前考虑的方案)

先将PDF转成带格式的HTML,再解析标签拆分:

  1. 用pdf2htmlEX转HTML(命令行工具,需提前安装)
pdf2htmlEX --zoom 1.5 法律合同.pdf
  1. 用Python的BeautifulSoup解析HTML并拆分
from bs4 import BeautifulSoup

def split_from_html(html_path):
    with open(html_path, 'r', encoding='utf-8') as f:
        soup = BeautifulSoup(f, 'html.parser')
    
    chapters = []
    current_title, current_content = None, []
    # 假设章节标题是h1标签,内容为后续的p/ul/li标签,可根据实际HTML结构调整
    for element in soup.find_all(['h1', 'p', 'ul', 'li']):
        if element.name == 'h1':
            if current_title:
                chapters.append((current_title, '\n'.join(current_content)))
            current_title = element.get_text(strip=True)
            current_content = []
        else:
            if current_title:
                text = element.get_text(strip=True)
                if text:
                    current_content.append(text)
    
    # 处理最后一个章节
    if current_title:
        chapters.append((current_title, '\n'.join(current_content)))
    
    # 写入文件
    for idx, (title, content) in enumerate(chapters, 1):
        filename = f"章节{idx}_{title}.txt"
        with open(filename, 'w', encoding='utf-8') as f:
            f.write(f"{title}\n\n{content}")
        print(f"已生成:{filename}")

# 替换为你的HTML路径
split_from_html("法律合同.html")

注意事项

  • 正则表达式和HTML标签选择器必须匹配你的合同实际格式,比如标题是“Article X”或“第X条”时,要对应修改规则
  • 如果PDF是扫描件(图片格式),需先用Tesseract做OCR识别文本,再用上述方法拆分

内容的提问来源于stack exchange,提问作者Akshat Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 18:57:20