如何将PDF法律合同的章节拆分至独立文本文件?
拆分PDF法律合同章节为独立文本文件的可行方案
方法一:Python脚本直接提取并拆分(推荐)
适合结构化清晰的PDF合同,通过正则匹配章节标题实现精准拆分:
- 安装依赖
pip install PyPDF2
- 运行以下脚本(根据合同标题格式调整正则)
import PyPDF2 import re def split_contract_chapters(pdf_path): # 读取PDF全量文本 with open(pdf_path, 'rb') as file: reader = PyPDF2.PdfReader(file) full_text = "" for page in reader.pages: full_text += page.extract_text() + "\n" # 匹配章节标题(示例:匹配“第X章 标题”格式,可按需修改正则) chapter_pattern = re.compile(r'(第\d+章\s+.+?)(?=第\d+章|$)', re.DOTALL) chapters = chapter_pattern.findall(full_text) # 生成独立文本文件 for idx, chapter in enumerate(chapters, 1): chapter_clean = chapter.strip() # 提取标题作为文件名 title_match = re.match(r'第\d+章\s+(.+)', chapter_clean.split('\n')[0]) filename = f"章节{idx}_{title_match.group(1)}.txt" if title_match else f"章节{idx}.txt" with open(filename, 'w', encoding='utf-8') as f: f.write(chapter_clean) print(f"已生成:{filename}") # 替换为你的PDF路径 split_contract_chapters("法律合同.pdf")
方法二:转HTML后解析拆分(你之前考虑的方案)
先将PDF转成带格式的HTML,再解析标签拆分:
- 用
pdf2htmlEX转HTML(命令行工具,需提前安装)
pdf2htmlEX --zoom 1.5 法律合同.pdf
- 用Python的BeautifulSoup解析HTML并拆分
from bs4 import BeautifulSoup def split_from_html(html_path): with open(html_path, 'r', encoding='utf-8') as f: soup = BeautifulSoup(f, 'html.parser') chapters = [] current_title, current_content = None, [] # 假设章节标题是h1标签,内容为后续的p/ul/li标签,可根据实际HTML结构调整 for element in soup.find_all(['h1', 'p', 'ul', 'li']): if element.name == 'h1': if current_title: chapters.append((current_title, '\n'.join(current_content))) current_title = element.get_text(strip=True) current_content = [] else: if current_title: text = element.get_text(strip=True) if text: current_content.append(text) # 处理最后一个章节 if current_title: chapters.append((current_title, '\n'.join(current_content))) # 写入文件 for idx, (title, content) in enumerate(chapters, 1): filename = f"章节{idx}_{title}.txt" with open(filename, 'w', encoding='utf-8') as f: f.write(f"{title}\n\n{content}") print(f"已生成:{filename}") # 替换为你的HTML路径 split_from_html("法律合同.html")
注意事项
- 正则表达式和HTML标签选择器必须匹配你的合同实际格式,比如标题是“Article X”或“第X条”时,要对应修改规则
- 如果PDF是扫描件(图片格式),需先用Tesseract做OCR识别文本,再用上述方法拆分
内容的提问来源于stack exchange,提问作者Akshat Gupta
相关产品推荐
相关产品推荐

