本地部署OCR管道文本输出优化:LLM返回前可靠清理乱码、控制字符及重复垃圾字符
I’ve tackled exactly this kind of OCR text cleanup in production-grade systems, and the following approach balances robustness, safety, and comprehensiveness to address all your requirements. It avoids the fragility of manual regex lists while protecting valid content like emails, URLs, and currency values.
1. Fix Mojibake & Encoding Errors Automatically
Manual replacement lists will always miss edge cases—ftfy is the gold standard here. Its fix_text method uses heuristic checks to detect and repair common encoding mishaps (like ’ instead of ’) without altering valid non-ASCII text. We wrap it in a try/except block to handle any unexpected edge cases gracefully:
from ftfy import fix_text def fix_mojibake(text: str) -> str: try: return fix_text(text) except Exception: # Fall back to original text if fix fails return text
2. Remove Invisible Control & Format Characters
OCR often inserts hidden control characters (zero-width spaces, RTL markers, soft hyphens) that break downstream processing. Two reliable methods here:
- Use
unicodedatato target all Unicode control categories (starts with "C") for full coverage - Use a regex for faster cleanup of common control ranges (ideal for high-volume pipelines)
Option A: Full Unicode Control Character Removal
import unicodedata def remove_all_control_chars(text: str) -> str: return "".join( ch for ch in text if not unicodedata.category(ch).startswith("C") )
Option B: Fast Regex-Based Cleanup
Targets the most common invisible characters found in OCR output:
import re def remove_common_control_chars(text: str) -> str: # Matches ASCII control chars, zero-width spaces, and format controls return re.sub(r'[\x00-\x1F\x7F-\x9F\u200B-\u200F\uFEFF]', '', text)
3. Normalize Compatibility & Ligature Characters
Unicode has many compatibility characters (like ligatures fl → fl, or variant currency symbols) that OCR often misinterprets. Use NFKC normalization to collapse these into their standard equivalents while preserving semantic meaning:
def normalize_compatibility_chars(text: str) -> str: return unicodedata.normalize("NFKC", text)
This will:
- Convert ligatures (
fi→fi,ffi→ffi) - Standardize currency symbols (e.g., variant rupee signs to
₹) - Fix superscript/subscript numbers to their base forms (use NFKD instead if you need to preserve formatting)
4. Collapse Excessive Repetitive Characters
OCR frequently generates long runs of repeated characters (e.g., aaaaaaaaaaaa...). We use a conservative regex to target only extreme repetitions (default: 50+ instances) and replace them with a concise marker to avoid altering valid content like repeated digits in serial numbers:
def collapse_excessive_repetitions(text: str, threshold: int = 50) -> str: # Match non-space characters repeated N+ times (with optional spaces in between) pattern = re.compile(r'(\S)(\s*\1){' + str(threshold-1) + r',}') def replacement(match): char = match.group(1) matched_text = match.group(0) # Handle spaced repetitions (e.g., "a a a ...") separately if ' ' in matched_text: return f"{char} {char} {char}..." # For continuous repetitions, use a tight marker return f"{char}{char}{char}..." return pattern.sub(replacement, text)
5. Preserve Critical Domain Tokens
To avoid accidentally modifying valid content like emails, URLs, or currency amounts, add a step to extract and reinsert these tokens after normalization. Here’s a simple example using regex matching:
def preserve_critical_tokens(text: str) -> str: # Define regex patterns for tokens to preserve token_patterns = [ r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', # Emails r'https?://\S+', # URLs r'\₹\d{1,3}(?:,\d{3})*(?:\.\d+)?', # Indian currency r'\b\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\b' # Timestamps ] # Extract tokens and replace with placeholders tokens = [] placeholder_template = f"__TOKEN_{{}}__" for pattern in token_patterns: for match in re.finditer(pattern, text): token = match.group(0) placeholder = placeholder_template.format(len(tokens)) text = text.replace(token, placeholder) tokens.append(token) # Apply core normalization steps normalized_text = normalize_ocr_text(text) # Reinsert preserved tokens for i, token in enumerate(tokens): normalized_text = normalized_text.replace(placeholder_template.format(i), token) return normalized_text
Full Integrated Pipeline
Combine all steps into a single function with a sensible order (fix encoding first, then normalize, clean controls, collapse repetitions):
def normalize_ocr_text(text: str) -> str: if not isinstance(text, str) or not text.strip(): return text # Step 1: Fix encoding/mojibake text = fix_mojibake(text) # Step 2: Normalize compatibility characters text = normalize_compatibility_chars(text) # Step 3: Remove control characters text = remove_all_control_chars(text) # Step 4: Collapse excessive repetitions text = collapse_excessive_repetitions(text) return text.strip()
Key Notes for Production:
- Test with your specific OCR dataset to adjust the repetition threshold (e.g., lower to 20 if your OCR generates shorter repeated runs)
- Add logging to track normalization changes for debugging
- For sensitive fields (like IDs), add field-specific validation to ensure no critical data is altered
内容的提问来源于stack exchange,提问作者agaonsindhe

