如何用Python正则表达式拆分含合同编号的Word文档?
实现方案(需求完全可行)
1. 核心思路
通过正则定位所有合同编号的起始位置,以此为拆分点将原文档文本分割为对应合同片段,最终保存为独立文件(优先选.txt,拆分/读取更高效;若需.docx可额外扩展)。
2. 分步实现代码
完整可运行代码
import os import re import docx2txt import textract def split_word_contracts(input_path, output_dir): # 创建输出目录(不存在则自动生成) os.makedirs(output_dir, exist_ok=True) # 读取不同格式的Word文档内容 if input_path.lower().endswith('.docx'): text = docx2txt.process(input_path) elif input_path.lower().endswith('.doc'): text_bytes = textract.process(input_path) text = text_bytes.decode('utf-8', errors='ignore') else: raise ValueError("仅支持.doc和.docx格式文件") # 正则匹配所有合同编号及其位置(兼容大小写、任意空白字符) pattern = re.compile(r'CONTRACT NUMBER\s+(.+)', re.IGNORECASE) matches = list(pattern.finditer(text)) if not matches: print("未找到任何合同编号") return # 拆分文本为单个合同片段 contract_fragments = [] for i in range(len(matches)): start_idx = matches[i].start() # 前n-1个合同:到下一个编号的起始位置结束 if i < len(matches) - 1: end_idx = matches[i+1].start() fragment = text[start_idx:end_idx].strip() # 最后一个合同:到文本结尾结束 else: fragment = text[start_idx:].strip() contract_fragments.append((matches[i].group(1).strip(), fragment)) # 保存每个合同文件 for _, (contract_num, content) in enumerate(contract_fragments, 1): # 替换编号中的非法字符,避免文件名报错 safe_num = re.sub(r'[^\w\-_.]', '_', contract_num) output_path = os.path.join(output_dir, f"合同_{safe_num}.txt") with open(output_path, 'w', encoding='utf-8') as f: f.write(content) print(f"已生成文件:{output_path}") # 调用示例(替换为你的文件路径) input_file = '/Users/aartimalik/Documents/GitHub/revenue_procurement/pdfs/bidsummaries-doc/100831R0.doc_2026.doc' output_directory = '/Users/aartimalik/Documents/GitHub/revenue_procurement/pdfs/bidsummaries-doc-test/split_contracts' split_word_contracts(input_file, output_directory)
3. 关键细节说明
- 格式兼容:
.docx用docx2txt轻量提取,.doc用textract(依赖系统工具如antiword,编码错误时用errors='ignore'避免崩溃)。 - 正则优化:加入
re.IGNORECASE处理大小写混合场景,\s+匹配任意空白字符(换行、多空格、制表符),覆盖更多合同编号格式。 - 文件名安全:替换编号中的非法字符(如
/、:),防止生成文件时出错。 - 扩展为.docx输出:若需生成Word格式文件,替换保存部分代码为:
from docx import Document doc = Document() doc.add_paragraph(content) doc.save(os.path.join(output_dir, f"合同_{safe_num}.docx"))
内容的提问来源于stack exchange,提问作者Pepa
相关产品推荐
相关产品推荐

