You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地化OCR Pipeline输出文本鲁棒清理方案:解决乱码(mojibake)、控制字符及重复字符问题

Robust Production-Grade OCR Text Normalization for On-Prem Pipelines

I’ve tackled exactly these kinds of OCR text cleanup issues in production systems before—manual regex and replacement lists quickly become unmaintainable, especially with the endless edge cases OCR throws at you. Here’s a battle-tested, Python-based solution that addresses all your requirements, with minimal overhead and maximum reliability:

Final Normalization Code (Ready to Integrate)

import re
import unicodedata
from ftfy import fix_text

def _remove_control_chars(s: str) -> str:
    # Remove ASCII control chars (0x00-0x1F) and extended control chars (0x7F-0x9F)
    # Covers zero-width spaces, soft hyphens, RTL markers, and other invisible garbage
    return re.sub(r'[\x00-\x1F\x7F-\x9F]', '', s)

def _collapse_repetitions(s: str, threshold: int = 50) -> str:
    # Collapse excessive non-space character repetitions (OCR blur/scanning artifacts)
    # Example: "aaaaaaaaaaaaaaaaaaaa..." → "aaa..."
    # Handles spaced repetitions too: "a a a a ..." → "a a a..."
    pattern = re.compile(r'(\S)(\s*\1){' + str(threshold-1) + r',}')
    
    def repl(match):
        char = match.group(1)
        full_match = match.group(0)
        if ' ' in full_match or '\n' in full_match:
            return f"{char} {char} {char}..."
        return char * 3 + "..."
    
    return pattern.sub(repl, s)

def _normalize_ocr_text(self, text: str) -> str:
    if not text or not isinstance(text, str):
        return text
    
    try:
        # Step 1: Auto-fix mojibake (encoding errors like ’ → ')
        text = fix_text(text)
    except Exception:
        # Fail-safe: if ftfy encounters an error, skip this step instead of breaking
        pass
    
    # Step 2: Normalize Unicode compatible characters/ligatures
    # Converts fl → fl, ₹ variants to standard ₹, etc.
    text = unicodedata.normalize("NFKC", text)
    
    # Step 3: Strip invisible control/formatting characters
    text = _remove_control_chars(text)
    
    # Step 4: Collapse excessive OCR-induced repetitions
    text = _collapse_repetitions(text)
    
    return text

Key Improvements Over Your Current Implementation

Let’s break down why this works better than manual rule-based cleanup:

  1. Auto-Mojibake Repair with ftfy

    • Replaces your manual mojibake_fixes dictionary with a library that uses heuristic encoding detection to fix hundreds of common encoding errors (e.g., UTF-8 misinterpreted as Latin-1, Windows-1252 artifacts).
    • Far more robust than manual lists—no need to update rules every time a new weird mojibake pops up.
  2. Unicode Normalization via NFKC

    • NFKC (Normalization Form Compatibility Composition) collapses:
      • Ligatures (e.g., ffi → ffi, fl → fl)
      • Compatibility characters (e.g., circled numbers → regular numbers, variant currency symbols → standard ones)
      • This ensures consistent text representation regardless of how OCR interpreted the original characters.
  3. Reliable Control Character Removal

    • The regex [\x00-\x1F\x7F-\x9F] targets all ASCII control characters and extended control/formatting characters (like zero-width spaces, soft hyphens, and bidirectional text markers) that OCR often introduces accidentally.
    • More straightforward than checking Unicode categories manually, while covering the same problematic characters.
  4. Conservative Repetition Collapsing

    • The threshold (default 50) ensures we only collapse truly excessive repetitions (OCR artifacts), not intentional repeated text.
    • Handles both contiguous repetitions (aaaaa...) and spaced repetitions (a a a...) with context-aware replacements.

Critical Edge Case: Preserving Domain-Specific Tokens

To avoid accidentally modifying critical tokens like emails, URLs, currency amounts, or timestamps, you can add a pre-check before applying repetition collapsing (or even normalization steps):

def _is_critical_token(s: str) -> bool:
    # Add regex patterns for your domain-specific tokens
    email_pattern = re.compile(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}')
    url_pattern = re.compile(r'https?://\S+')
    currency_pattern = re.compile(r'₹\d{1,3}(,\d{3})*(\.\d{2})?')
    
    return bool(email_pattern.search(s) or url_pattern.search(s) or currency_pattern.search(s))

# Then in _normalize_ocr_text:
if not _is_critical_token(text):
    text = _collapse_repetitions(text)

Integration with Your Existing Pipeline

Drop the _normalize_ocr_text method into your existing class, and it’ll work seamlessly with your current LLM call flow:

result = client.chat.completions.create(
    model=self.image_llm_model,
    messages=messages,
    max_tokens=max_tokens,
    response_format={"type": "json_object"},
)
parsed = json.loads(result.choices[0].message.content)
info_text = parsed.get("info")
return self._normalize_ocr_text(info_text) if info_text else info_text

内容的提问来源于stack exchange,提问作者agaonsindhe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 06:42:35