You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Regex提取多组合同字段的全部匹配结果

Solution for Extracting Multiple Contract Entries from Word Documents

Core Fix for Multiple Matches

  • Replace regex methods that only return the first match (like search()) with finditer() or findall() to capture all contract groups.
  • Design regex patterns to target complete contract blocks instead of isolated fields, ensuring each group of related variables is captured as a single entry.

Sample Implementation (Python)

Assuming you've extracted text from your Word document into a string doc_text:

Step 1: Define Block-Focused Regex Pattern

Adjust this pattern to match the exact structure of your contract text:

import re
import pandas as pd

# Regex to capture full contract blocks (tweak lookaheads/quantifiers for your text)
contract_pattern = re.compile(
    r"Contract Number:\s*(\S+)\s*"
    r"Location:\s*(.+?)\s*(?=Contract Items|Contract Code|Federal Aid|$)\s*"
    r"Contract Items:\s*(.+?)\s*(?=Contract Code|Federal Aid|$)\s*"
    r"Contract Code:\s*(\S+)\s*"
    r"Federal Aid:\s*(\S+)",
    re.DOTALL | re.IGNORECASE
)

# Extract all matching contract blocks
matches = contract_pattern.finditer(doc_text)

# Convert matches to structured data
contract_list = []
for match in matches:
    contract_list.append({
        "contract_number": match.group(1),
        "location": match.group(2).strip(),
        "contract_items": match.group(3).strip(),
        "contract_code": match.group(4),
        "federal_aid": match.group(5)
    })

# Generate tabular output
df = pd.DataFrame(contract_list)
print(df)

Pattern Adjustment Notes

  • Use re.DOTALL to let . match newlines (critical for multi-line contract blocks).
  • Non-greedy quantifiers (+?) prevent the pattern from spanning across multiple contract entries.
  • Update lookaheads ((?=...)) to match the actual delimiters between contracts in your text.

Loop Over Multiple Documents

Wrap the extraction logic in a function to process batches of Word files:

from docx import Document
import os

def extract_contracts(doc_path):
    # Read Word document text
    doc = Document(doc_path)
    doc_text = "\n".join([para.text for para in doc.paragraphs])
    
    # Extract contracts
    matches = contract_pattern.finditer(doc_text)
    entries = []
    for match in matches:
        entries.append({
            "source_doc": doc_path,
            "contract_number": match.group(1),
            "location": match.group(2).strip(),
            "contract_items": match.group(3).strip(),
            "contract_code": match.group(4),
            "federal_aid": match.group(5)
        })
    return entries

# Process all .docx files in a directory
all_entries = []
target_dir = "/path/to/your/contract_docs"

for filename in os.listdir(target_dir):
    if filename.endswith(".docx"):
        full_path = os.path.join(target_dir, filename)
        all_entries.extend(extract_contracts(full_path))

# Save combined results to CSV
final_df = pd.DataFrame(all_entries)
final_df.to_csv("all_contracts.csv", index=False)

Expected Output Example

source_doccontract_numberlocationcontract_itemscontract_codefederal_aid
contracts_sample.docCN-001New YorkOffice Supplies, FurnitureCC-101YES
contracts_sample.docCN-002Los AngelesIT EquipmentCC-102NO
contracts_sample.docCN-003ChicagoBuilding MaterialsCC-103YES

Key Tips

  • Test your regex pattern against sample text first to refine capturing groups.
  • Add optional groups (e.g., (?:Federal Aid:\s*(\S+))?) to handle missing fields gracefully.
  • Install required libraries with pip install python-docx pandas.

内容的提问来源于stack exchange,提问作者Pepa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 11:05:20