本地化OCR Pipeline输出文本鲁棒清理方案:解决乱码(mojibake)、控制字符及重复字符问题
I’ve tackled exactly these kinds of OCR text cleanup issues in production systems before—manual regex and replacement lists quickly become unmaintainable, especially with the endless edge cases OCR throws at you. Here’s a battle-tested, Python-based solution that addresses all your requirements, with minimal overhead and maximum reliability:
Final Normalization Code (Ready to Integrate)
import re import unicodedata from ftfy import fix_text def _remove_control_chars(s: str) -> str: # Remove ASCII control chars (0x00-0x1F) and extended control chars (0x7F-0x9F) # Covers zero-width spaces, soft hyphens, RTL markers, and other invisible garbage return re.sub(r'[\x00-\x1F\x7F-\x9F]', '', s) def _collapse_repetitions(s: str, threshold: int = 50) -> str: # Collapse excessive non-space character repetitions (OCR blur/scanning artifacts) # Example: "aaaaaaaaaaaaaaaaaaaa..." → "aaa..." # Handles spaced repetitions too: "a a a a ..." → "a a a..." pattern = re.compile(r'(\S)(\s*\1){' + str(threshold-1) + r',}') def repl(match): char = match.group(1) full_match = match.group(0) if ' ' in full_match or '\n' in full_match: return f"{char} {char} {char}..." return char * 3 + "..." return pattern.sub(repl, s) def _normalize_ocr_text(self, text: str) -> str: if not text or not isinstance(text, str): return text try: # Step 1: Auto-fix mojibake (encoding errors like ’ → ') text = fix_text(text) except Exception: # Fail-safe: if ftfy encounters an error, skip this step instead of breaking pass # Step 2: Normalize Unicode compatible characters/ligatures # Converts fl → fl, ₹ variants to standard ₹, etc. text = unicodedata.normalize("NFKC", text) # Step 3: Strip invisible control/formatting characters text = _remove_control_chars(text) # Step 4: Collapse excessive OCR-induced repetitions text = _collapse_repetitions(text) return text
Key Improvements Over Your Current Implementation
Let’s break down why this works better than manual rule-based cleanup:
Auto-Mojibake Repair with
ftfy- Replaces your manual
mojibake_fixesdictionary with a library that uses heuristic encoding detection to fix hundreds of common encoding errors (e.g., UTF-8 misinterpreted as Latin-1, Windows-1252 artifacts). - Far more robust than manual lists—no need to update rules every time a new weird mojibake pops up.
- Replaces your manual
Unicode Normalization via NFKC
NFKC(Normalization Form Compatibility Composition) collapses:- Ligatures (e.g.,
ffi→ffi,fl→fl) - Compatibility characters (e.g., circled numbers → regular numbers, variant currency symbols → standard ones)
- This ensures consistent text representation regardless of how OCR interpreted the original characters.
- Ligatures (e.g.,
Reliable Control Character Removal
- The regex
[\x00-\x1F\x7F-\x9F]targets all ASCII control characters and extended control/formatting characters (like zero-width spaces, soft hyphens, and bidirectional text markers) that OCR often introduces accidentally. - More straightforward than checking Unicode categories manually, while covering the same problematic characters.
- The regex
Conservative Repetition Collapsing
- The threshold (default 50) ensures we only collapse truly excessive repetitions (OCR artifacts), not intentional repeated text.
- Handles both contiguous repetitions (
aaaaa...) and spaced repetitions (a a a...) with context-aware replacements.
Critical Edge Case: Preserving Domain-Specific Tokens
To avoid accidentally modifying critical tokens like emails, URLs, currency amounts, or timestamps, you can add a pre-check before applying repetition collapsing (or even normalization steps):
def _is_critical_token(s: str) -> bool: # Add regex patterns for your domain-specific tokens email_pattern = re.compile(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}') url_pattern = re.compile(r'https?://\S+') currency_pattern = re.compile(r'₹\d{1,3}(,\d{3})*(\.\d{2})?') return bool(email_pattern.search(s) or url_pattern.search(s) or currency_pattern.search(s)) # Then in _normalize_ocr_text: if not _is_critical_token(text): text = _collapse_repetitions(text)
Integration with Your Existing Pipeline
Drop the _normalize_ocr_text method into your existing class, and it’ll work seamlessly with your current LLM call flow:
result = client.chat.completions.create( model=self.image_llm_model, messages=messages, max_tokens=max_tokens, response_format={"type": "json_object"}, ) parsed = json.loads(result.choices[0].message.content) info_text = parsed.get("info") return self._normalize_ocr_text(info_text) if info_text else info_text
内容的提问来源于stack exchange,提问作者agaonsindhe

