You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式拆分含合同编号的Word文档?

实现方案(需求完全可行)

1. 核心思路

通过正则定位所有合同编号的起始位置,以此为拆分点将原文档文本分割为对应合同片段,最终保存为独立文件(优先选.txt,拆分/读取更高效;若需.docx可额外扩展)。

2. 分步实现代码

完整可运行代码

import os
import re
import docx2txt
import textract

def split_word_contracts(input_path, output_dir):
    # 创建输出目录(不存在则自动生成)
    os.makedirs(output_dir, exist_ok=True)
    
    # 读取不同格式的Word文档内容
    if input_path.lower().endswith('.docx'):
        text = docx2txt.process(input_path)
    elif input_path.lower().endswith('.doc'):
        text_bytes = textract.process(input_path)
        text = text_bytes.decode('utf-8', errors='ignore')
    else:
        raise ValueError("仅支持.doc和.docx格式文件")
    
    # 正则匹配所有合同编号及其位置(兼容大小写、任意空白字符)
    pattern = re.compile(r'CONTRACT NUMBER\s+(.+)', re.IGNORECASE)
    matches = list(pattern.finditer(text))
    
    if not matches:
        print("未找到任何合同编号")
        return
    
    # 拆分文本为单个合同片段
    contract_fragments = []
    for i in range(len(matches)):
        start_idx = matches[i].start()
        # 前n-1个合同:到下一个编号的起始位置结束
        if i < len(matches) - 1:
            end_idx = matches[i+1].start()
            fragment = text[start_idx:end_idx].strip()
        # 最后一个合同:到文本结尾结束
        else:
            fragment = text[start_idx:].strip()
        contract_fragments.append((matches[i].group(1).strip(), fragment))
    
    # 保存每个合同文件
    for _, (contract_num, content) in enumerate(contract_fragments, 1):
        # 替换编号中的非法字符,避免文件名报错
        safe_num = re.sub(r'[^\w\-_.]', '_', contract_num)
        output_path = os.path.join(output_dir, f"合同_{safe_num}.txt")
        with open(output_path, 'w', encoding='utf-8') as f:
            f.write(content)
        print(f"已生成文件:{output_path}")

# 调用示例(替换为你的文件路径)
input_file = '/Users/aartimalik/Documents/GitHub/revenue_procurement/pdfs/bidsummaries-doc/100831R0.doc_2026.doc'
output_directory = '/Users/aartimalik/Documents/GitHub/revenue_procurement/pdfs/bidsummaries-doc-test/split_contracts'
split_word_contracts(input_file, output_directory)

3. 关键细节说明

  • 格式兼容:.docx用docx2txt轻量提取,.doc用textract(依赖系统工具如antiword,编码错误时用errors='ignore'避免崩溃)。
  • 正则优化:加入re.IGNORECASE处理大小写混合场景,\s+匹配任意空白字符(换行、多空格、制表符),覆盖更多合同编号格式。
  • 文件名安全:替换编号中的非法字符(如/、:),防止生成文件时出错。
  • 扩展为.docx输出:若需生成Word格式文件,替换保存部分代码为:
    from docx import Document
    doc = Document()
    doc.add_paragraph(content)
    doc.save(os.path.join(output_dir, f"合同_{safe_num}.docx"))
    

内容的提问来源于stack exchange,提问作者Pepa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 01:00:58