You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP中非等长文本比对咨询:DOCX与TXT动态内容匹配方案

Robust Solution for Comparing DOCX Templates and Filled TXT Content

Your core issue is relying on hardcoded split points that break when the fixed text in your template changes. Instead, we can leverage the template's placeholder markers to dynamically identify fixed vs variable sections, enabling reliable comparison even for long documents.

Approach Overview

  1. Parse the DOCX Template: Extract fixed text segments and placeholder positions using regex.
  2. Validate Filled Content Structure: Use a dynamically generated regex pattern to check if the TXT file follows the template's structure.
  3. Compare Fixed Segments: Directly compare the fixed text parts between the template and filled content (ignoring placeholder values) to detect any changes.

Step-by-Step Implementation

1. Install Dependencies

First, install python-docx to read DOCX files:

pip install python-docx

2. Code to Parse the DOCX Template

This code reads the DOCX and splits it into fixed segments and placeholders:

import re
from docx import Document

def read_docx_template(docx_path):
    """Read all text from a DOCX file, preserving paragraph structure."""
    doc = Document(docx_path)
    return '\n'.join([para.text for para in doc.paragraphs])

# Load your template (replace with your file path)
template_text = read_docx_template("template.docx")

# Split template into alternating fixed text and placeholders
# Regex matches {PLACEHOLDER} patterns
parts = re.split(r'(\{.*?\})', template_text)

# Separate fixed segments (non-placeholder text) and placeholders
fixed_segments = [
    part.strip() for part in parts 
    if not re.match(r'\{.*?\}', part) and part.strip()  # Exclude empty strings
]
placeholders = [part for part in parts if re.match(r'\{.*?\}', part)]

print("Template Fixed Segments:", fixed_segments)
print("Template Placeholders:", placeholders)

3. Validate Filled TXT Structure and Extract Placeholders

This code checks if the TXT file conforms to the template's structure and can extract filled placeholder values:

# Read filled TXT content (replace with your file path)
with open("filled.txt", "r") as f:
    filled_text = f.read()

# Build a regex pattern from the template
pattern_components = []
for part in parts:
    if re.match(r'\{.*?\}', part):
        # Replace placeholders with non-greedy wildcards to match any content
        pattern_components.append(r'(.*?)')
    else:
        # Escape fixed text to handle regex special characters (like dots, commas)
        pattern_components.append(re.escape(part))

template_regex = ''.join(pattern_components)

# Check if filled text matches the template structure (allow newlines with re.DOTALL)
match = re.fullmatch(template_regex, filled_text, re.DOTALL)
if match:
    print("\nFilled content matches template structure!")
    # Extract and print placeholder values
    placeholder_values = match.groups()
    for placeholder, value in zip(placeholders, placeholder_values):
        print(f"{placeholder}: {value.strip()}")
else:
    print("\nFilled content does NOT match template structure.")

4. Compare Fixed Segments for Exact Matches

This function checks if all fixed text segments from the template are present (and unchanged) in the filled TXT, ignoring placeholder content:

def compare_fixed_content(template_fixed, filled_text):
    """Compare fixed segments between template and filled content, normalizing whitespace."""
    current_position = 0
    for idx, segment in enumerate(template_fixed):
        # Find the segment in the filled text starting from current position
        segment_pos = filled_text.find(segment, current_position)
        if segment_pos == -1:
            print(f"\nMismatch at segment {idx+1}: Missing segment")
            print(f"Template segment: {segment}")
            return False
        
        # Extract the corresponding segment from filled text
        filled_segment = filled_text[segment_pos : segment_pos + len(segment)]
        
        # Normalize whitespace to handle minor formatting differences
        template_normalized = ' '.join(segment.split())
        filled_normalized = ' '.join(filled_segment.split())
        
        if template_normalized != filled_normalized:
            print(f"\nMismatch at segment {idx+1}:")
            print(f"Template: {segment}")
            print(f"Filled: {filled_segment}")
            return False
        
        current_position = segment_pos + len(segment)
    
    print("\nAll fixed segments match exactly!")
    return True

# Run the comparison
compare_fixed_content(fixed_segments, filled_text)

Key Advantages

  • Dynamic Structure Detection: Doesn't rely on hardcoded split strings—uses the template's own placeholders to guide comparison.
  • Whitespace Tolerance: Normalizes whitespace to handle minor formatting differences (e.g., extra spaces, line breaks).
  • Scalable for Long Documents: Works seamlessly with 20-page DOCX files since it processes text linearly.
  • Change Detection: Automatically flags any changes to fixed text segments, which is critical for verifying template consistency.

内容的提问来源于stack exchange,提问作者mr Pi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:43:28