You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地部署OCR管道文本输出优化:LLM返回前可靠清理乱码、控制字符及重复垃圾字符

Robust OCR Text Normalization Pipeline for Production

I’ve tackled exactly this kind of OCR text cleanup in production-grade systems, and the following approach balances robustness, safety, and comprehensiveness to address all your requirements. It avoids the fragility of manual regex lists while protecting valid content like emails, URLs, and currency values.

1. Fix Mojibake & Encoding Errors Automatically

Manual replacement lists will always miss edge cases—ftfy is the gold standard here. Its fix_text method uses heuristic checks to detect and repair common encoding mishaps (like ’ instead of ’) without altering valid non-ASCII text. We wrap it in a try/except block to handle any unexpected edge cases gracefully:

from ftfy import fix_text

def fix_mojibake(text: str) -> str:
    try:
        return fix_text(text)
    except Exception:
        # Fall back to original text if fix fails
        return text

2. Remove Invisible Control & Format Characters

OCR often inserts hidden control characters (zero-width spaces, RTL markers, soft hyphens) that break downstream processing. Two reliable methods here:

  • Use unicodedata to target all Unicode control categories (starts with "C") for full coverage
  • Use a regex for faster cleanup of common control ranges (ideal for high-volume pipelines)

Option A: Full Unicode Control Character Removal

import unicodedata

def remove_all_control_chars(text: str) -> str:
    return "".join(
        ch for ch in text 
        if not unicodedata.category(ch).startswith("C")
    )

Option B: Fast Regex-Based Cleanup

Targets the most common invisible characters found in OCR output:

import re

def remove_common_control_chars(text: str) -> str:
    # Matches ASCII control chars, zero-width spaces, and format controls
    return re.sub(r'[\x00-\x1F\x7F-\x9F\u200B-\u200F\uFEFF]', '', text)

3. Normalize Compatibility & Ligature Characters

Unicode has many compatibility characters (like ligatures fl → fl, or variant currency symbols) that OCR often misinterprets. Use NFKC normalization to collapse these into their standard equivalents while preserving semantic meaning:

def normalize_compatibility_chars(text: str) -> str:
    return unicodedata.normalize("NFKC", text)

This will:

  • Convert ligatures (fi → fi, ffi → ffi)
  • Standardize currency symbols (e.g., variant rupee signs to ₹)
  • Fix superscript/subscript numbers to their base forms (use NFKD instead if you need to preserve formatting)

4. Collapse Excessive Repetitive Characters

OCR frequently generates long runs of repeated characters (e.g., aaaaaaaaaaaa...). We use a conservative regex to target only extreme repetitions (default: 50+ instances) and replace them with a concise marker to avoid altering valid content like repeated digits in serial numbers:

def collapse_excessive_repetitions(text: str, threshold: int = 50) -> str:
    # Match non-space characters repeated N+ times (with optional spaces in between)
    pattern = re.compile(r'(\S)(\s*\1){' + str(threshold-1) + r',}')
    
    def replacement(match):
        char = match.group(1)
        matched_text = match.group(0)
        # Handle spaced repetitions (e.g., "a a a ...") separately
        if ' ' in matched_text:
            return f"{char} {char} {char}..."
        # For continuous repetitions, use a tight marker
        return f"{char}{char}{char}..."
    
    return pattern.sub(replacement, text)

5. Preserve Critical Domain Tokens

To avoid accidentally modifying valid content like emails, URLs, or currency amounts, add a step to extract and reinsert these tokens after normalization. Here’s a simple example using regex matching:

def preserve_critical_tokens(text: str) -> str:
    # Define regex patterns for tokens to preserve
    token_patterns = [
        r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b',  # Emails
        r'https?://\S+',  # URLs
        r'\₹\d{1,3}(?:,\d{3})*(?:\.\d+)?',  # Indian currency
        r'\b\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\b'  # Timestamps
    ]
    
    # Extract tokens and replace with placeholders
    tokens = []
    placeholder_template = f"__TOKEN_{{}}__"
    
    for pattern in token_patterns:
        for match in re.finditer(pattern, text):
            token = match.group(0)
            placeholder = placeholder_template.format(len(tokens))
            text = text.replace(token, placeholder)
            tokens.append(token)
    
    # Apply core normalization steps
    normalized_text = normalize_ocr_text(text)
    
    # Reinsert preserved tokens
    for i, token in enumerate(tokens):
        normalized_text = normalized_text.replace(placeholder_template.format(i), token)
    
    return normalized_text

Full Integrated Pipeline

Combine all steps into a single function with a sensible order (fix encoding first, then normalize, clean controls, collapse repetitions):

def normalize_ocr_text(text: str) -> str:
    if not isinstance(text, str) or not text.strip():
        return text
    
    # Step 1: Fix encoding/mojibake
    text = fix_mojibake(text)
    # Step 2: Normalize compatibility characters
    text = normalize_compatibility_chars(text)
    # Step 3: Remove control characters
    text = remove_all_control_chars(text)
    # Step 4: Collapse excessive repetitions
    text = collapse_excessive_repetitions(text)
    
    return text.strip()

Key Notes for Production:

  • Test with your specific OCR dataset to adjust the repetition threshold (e.g., lower to 20 if your OCR generates shorter repeated runs)
  • Add logging to track normalization changes for debugging
  • For sensitive fields (like IDs), add field-specific validation to ensure no critical data is altered

内容的提问来源于stack exchange,提问作者agaonsindhe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 06:40:13