如何用Python Regex提取多组合同字段的全部匹配结果
Solution for Extracting Multiple Contract Entries from Word Documents
Core Fix for Multiple Matches
- Replace regex methods that only return the first match (like
search()) withfinditer()orfindall()to capture all contract groups. - Design regex patterns to target complete contract blocks instead of isolated fields, ensuring each group of related variables is captured as a single entry.
Sample Implementation (Python)
Assuming you've extracted text from your Word document into a string doc_text:
Step 1: Define Block-Focused Regex Pattern
Adjust this pattern to match the exact structure of your contract text:
import re import pandas as pd # Regex to capture full contract blocks (tweak lookaheads/quantifiers for your text) contract_pattern = re.compile( r"Contract Number:\s*(\S+)\s*" r"Location:\s*(.+?)\s*(?=Contract Items|Contract Code|Federal Aid|$)\s*" r"Contract Items:\s*(.+?)\s*(?=Contract Code|Federal Aid|$)\s*" r"Contract Code:\s*(\S+)\s*" r"Federal Aid:\s*(\S+)", re.DOTALL | re.IGNORECASE ) # Extract all matching contract blocks matches = contract_pattern.finditer(doc_text) # Convert matches to structured data contract_list = [] for match in matches: contract_list.append({ "contract_number": match.group(1), "location": match.group(2).strip(), "contract_items": match.group(3).strip(), "contract_code": match.group(4), "federal_aid": match.group(5) }) # Generate tabular output df = pd.DataFrame(contract_list) print(df)
Pattern Adjustment Notes
- Use
re.DOTALLto let.match newlines (critical for multi-line contract blocks). - Non-greedy quantifiers (
+?) prevent the pattern from spanning across multiple contract entries. - Update lookaheads (
(?=...)) to match the actual delimiters between contracts in your text.
Loop Over Multiple Documents
Wrap the extraction logic in a function to process batches of Word files:
from docx import Document import os def extract_contracts(doc_path): # Read Word document text doc = Document(doc_path) doc_text = "\n".join([para.text for para in doc.paragraphs]) # Extract contracts matches = contract_pattern.finditer(doc_text) entries = [] for match in matches: entries.append({ "source_doc": doc_path, "contract_number": match.group(1), "location": match.group(2).strip(), "contract_items": match.group(3).strip(), "contract_code": match.group(4), "federal_aid": match.group(5) }) return entries # Process all .docx files in a directory all_entries = [] target_dir = "/path/to/your/contract_docs" for filename in os.listdir(target_dir): if filename.endswith(".docx"): full_path = os.path.join(target_dir, filename) all_entries.extend(extract_contracts(full_path)) # Save combined results to CSV final_df = pd.DataFrame(all_entries) final_df.to_csv("all_contracts.csv", index=False)
Expected Output Example
| source_doc | contract_number | location | contract_items | contract_code | federal_aid |
|---|---|---|---|---|---|
| contracts_sample.doc | CN-001 | New York | Office Supplies, Furniture | CC-101 | YES |
| contracts_sample.doc | CN-002 | Los Angeles | IT Equipment | CC-102 | NO |
| contracts_sample.doc | CN-003 | Chicago | Building Materials | CC-103 | YES |
Key Tips
- Test your regex pattern against sample text first to refine capturing groups.
- Add optional groups (e.g.,
(?:Federal Aid:\s*(\S+))?) to handle missing fields gracefully. - Install required libraries with
pip install python-docx pandas.
内容的提问来源于stack exchange,提问作者Pepa
相关产品推荐
相关产品推荐

