You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用spaCy规范化格式化文本以消除年度财报对比中的无意义差异?

Solution: Standardize Text Formatting for Meaningful Diff Comparisons

Great question—this is a common pain point when working with converted PDF text. Let’s break down how to use spaCy and basic text processing to standardize your documents so diff only flags real content changes, not formatting artifacts.

Key Steps to Fix Your Problem

We’ll address both the inconsistent bullets and line break issues, plus handle minor content shifts gracefully:

  1. Normalize Bullet Points: Replace all non-standard bullets (like x or ●) with a uniform marker (e.g., -).
  2. Clean & Segment Text: Use spaCy to split content into sentences—this lets us isolate minor differences without breaking the entire document’s formatting.
  3. Consistent Line Wrapping: Wrap each sentence into lines of max 80 characters, breaking only at word boundaries to avoid false diffs.
  4. Handle Minor Content Shifts: By working at the sentence level, line breaks only shift where actual content differs, keeping the rest of the text aligned.

Full Implementation Code

Here’s a runnable script that applies these steps to your sample texts:

import spacy
import re

# Load spaCy's English model
nlp = spacy.load("en_core_web_sm")

def normalize_bullets_and_split_items(text):
    """Split text into list items and replace inconsistent bullets with "-"."""
    # Split on any bullet-like symbols (x, ●) that are at the start or after a semicolon
    items = re.split(r'\s*(x|●)\s+', text.strip())
    # Filter out empty strings and bullet symbols, clean trailing semicolons
    cleaned_items = []
    for i in range(1, len(items), 2):
        item_content = items[i+1].strip().rstrip(';')
        if item_content:
            cleaned_items.append(item_content)
    # Add standard bullet to each item
    return [f"- {item}" for item in cleaned_items]

def wrap_sentence(sentence_text, max_line_length=80):
    """Wrap a single sentence into lines of max N characters, breaking at word boundaries."""
    words = sentence_text.split()
    lines = []
    current_line = []
    current_length = 0
    
    for word in words:
        # Calculate length if we add this word (account for space between words)
        next_length = current_length + len(word) + (1 if current_length > 0 else 0)
        if next_length <= max_line_length or current_length == 0:
            current_line.append(word)
            current_length = next_length
        else:
            lines.append(' '.join(current_line))
            current_line = [word]
            current_length = len(word)
    
    if current_line:
        lines.append(' '.join(current_line))
    return lines

def format_document(text):
    """Full pipeline to convert raw text into standardized, diff-friendly format."""
    # Step 1: Normalize bullets and split into list items
    list_items = normalize_bullets_and_split_items(text)
    formatted_lines = []
    
    for item in list_items:
        # Step 2: Use spaCy to split the item into sentences
        doc = nlp(item[2:])  # Skip the "- " bullet for processing
        sentence_lines = []
        
        for sent in doc.sents:
            # Reconstruct clean sentence text (preserves proper spacing)
            sent_text = sent.text.strip()
            # Step 3: Wrap each sentence into lines
            wrapped = wrap_sentence(sent_text)
            # Indent subsequent lines of the same sentence for readability
            sentence_lines.append(wrapped[0])
            sentence_lines.extend([f"  {line}" for line in wrapped[1:]])
        
        # Combine lines for the list item, starting with the bullet
        formatted_lines.append(f"- {sentence_lines[0]}")
        formatted_lines.extend(sentence_lines[1:])
    
    return '\n'.join(formatted_lines)

# Test with your sample texts
SE_2018_10k_string = '''x “paying users” refers to the number of unique accounts through which a payment is made in our online games in a particular period. A unique account through which payments are made in more than one online game or in more than one market is counted as more than one paying user. “QPUs” refers to the aggregate number of paying users during the quarterly period; x'''
SE_2019_10k_string = '''● “paying users” refers to the number of unique accounts through which a payment is made in our online games in a particular period. A unique account through which payments are made in more than one online game or in more than one market is counted as more than one paying user. “QPUs” refers to the aggregate number of paying users during the quarterly period; ●'''

# Format both documents
formatted_2018 = format_document(SE_2018_10k_string)
formatted_2019 = format_document(SE_2019_10k_string)

# Print results
print("Formatted 2018 Text:\n")
print(formatted_2018)
print("\n---\n")
print("Formatted 2019 Text:\n")
print(formatted_2019)

# Check similarity (should be very high since content is identical)
doc1 = nlp(SE_2018_10k_string)
doc2 = nlp(SE_2019_10k_string)
print(f"\nDocument Similarity: {doc1.similarity(doc2):.4f}")

What This Does

  • Bullet Normalization: The normalize_bullets_and_split_items function uses regex to split your text into individual list items and replaces all inconsistent bullets with -.
  • Sentence Segmentation: SpaCy’s sentencizer splits each list item into sentences, so minor content changes (like an extra word) only affect the lines for that specific sentence, not the entire document.
  • Line Wrapping: The wrap_sentence function builds lines word-by-word to ensure we never break mid-word, and keeps lines under 80 characters. Subsequent lines of the same sentence are indented for readability.

Example Output

Both formatted texts will be identical (since your samples have no content differences), so diff will show no changes. If one version had an extra word (e.g., "A unique account that through payments..."), only the lines for that sentence would differ in the diff output.

Handling Edge Cases

  • Extra Punctuation: The code strips trailing semicolons from list items (common PDF conversion artifacts).
  • Long Words: If a word is longer than 80 characters (unlikely in financial reports), the function will still include it as a single line to avoid breaking it.
  • Multiple List Items: The code handles multiple list items correctly, splitting them and applying the same formatting rules to each.

内容的提问来源于stack exchange,提问作者daeda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 09:07:29