NLP中非等长文本比对咨询:DOCX与TXT动态内容匹配方案
Robust Solution for Comparing DOCX Templates and Filled TXT Content
Your core issue is relying on hardcoded split points that break when the fixed text in your template changes. Instead, we can leverage the template's placeholder markers to dynamically identify fixed vs variable sections, enabling reliable comparison even for long documents.
Approach Overview
- Parse the DOCX Template: Extract fixed text segments and placeholder positions using regex.
- Validate Filled Content Structure: Use a dynamically generated regex pattern to check if the TXT file follows the template's structure.
- Compare Fixed Segments: Directly compare the fixed text parts between the template and filled content (ignoring placeholder values) to detect any changes.
Step-by-Step Implementation
1. Install Dependencies
First, install python-docx to read DOCX files:
pip install python-docx
2. Code to Parse the DOCX Template
This code reads the DOCX and splits it into fixed segments and placeholders:
import re from docx import Document def read_docx_template(docx_path): """Read all text from a DOCX file, preserving paragraph structure.""" doc = Document(docx_path) return '\n'.join([para.text for para in doc.paragraphs]) # Load your template (replace with your file path) template_text = read_docx_template("template.docx") # Split template into alternating fixed text and placeholders # Regex matches {PLACEHOLDER} patterns parts = re.split(r'(\{.*?\})', template_text) # Separate fixed segments (non-placeholder text) and placeholders fixed_segments = [ part.strip() for part in parts if not re.match(r'\{.*?\}', part) and part.strip() # Exclude empty strings ] placeholders = [part for part in parts if re.match(r'\{.*?\}', part)] print("Template Fixed Segments:", fixed_segments) print("Template Placeholders:", placeholders)
3. Validate Filled TXT Structure and Extract Placeholders
This code checks if the TXT file conforms to the template's structure and can extract filled placeholder values:
# Read filled TXT content (replace with your file path) with open("filled.txt", "r") as f: filled_text = f.read() # Build a regex pattern from the template pattern_components = [] for part in parts: if re.match(r'\{.*?\}', part): # Replace placeholders with non-greedy wildcards to match any content pattern_components.append(r'(.*?)') else: # Escape fixed text to handle regex special characters (like dots, commas) pattern_components.append(re.escape(part)) template_regex = ''.join(pattern_components) # Check if filled text matches the template structure (allow newlines with re.DOTALL) match = re.fullmatch(template_regex, filled_text, re.DOTALL) if match: print("\nFilled content matches template structure!") # Extract and print placeholder values placeholder_values = match.groups() for placeholder, value in zip(placeholders, placeholder_values): print(f"{placeholder}: {value.strip()}") else: print("\nFilled content does NOT match template structure.")
4. Compare Fixed Segments for Exact Matches
This function checks if all fixed text segments from the template are present (and unchanged) in the filled TXT, ignoring placeholder content:
def compare_fixed_content(template_fixed, filled_text): """Compare fixed segments between template and filled content, normalizing whitespace.""" current_position = 0 for idx, segment in enumerate(template_fixed): # Find the segment in the filled text starting from current position segment_pos = filled_text.find(segment, current_position) if segment_pos == -1: print(f"\nMismatch at segment {idx+1}: Missing segment") print(f"Template segment: {segment}") return False # Extract the corresponding segment from filled text filled_segment = filled_text[segment_pos : segment_pos + len(segment)] # Normalize whitespace to handle minor formatting differences template_normalized = ' '.join(segment.split()) filled_normalized = ' '.join(filled_segment.split()) if template_normalized != filled_normalized: print(f"\nMismatch at segment {idx+1}:") print(f"Template: {segment}") print(f"Filled: {filled_segment}") return False current_position = segment_pos + len(segment) print("\nAll fixed segments match exactly!") return True # Run the comparison compare_fixed_content(fixed_segments, filled_text)
Key Advantages
- Dynamic Structure Detection: Doesn't rely on hardcoded split strings—uses the template's own placeholders to guide comparison.
- Whitespace Tolerance: Normalizes whitespace to handle minor formatting differences (e.g., extra spaces, line breaks).
- Scalable for Long Documents: Works seamlessly with 20-page DOCX files since it processes text linearly.
- Change Detection: Automatically flags any changes to fixed text segments, which is critical for verifying template consistency.
内容的提问来源于stack exchange,提问作者mr Pi
相关产品推荐
相关产品推荐

