求助:OCR生成文本文件的拼写检查与纠错效率优化
Hey there! Let's tackle your OCR text spell-checking efficiency issues step by step. OCR-generated text has unique error patterns (like character confusions: i/l/1, o/0, s/5, etc.) that generic spell-checkers often struggle with, so tweaking your approach for these specifics will make a huge difference. Here are actionable optimizations tailored to your use case:
1. Optimize for OCR-Specific Error Patterns
OCR errors aren't just typos—they're often character misrecognitions. Addressing these first reduces the load on your spell-checker:
- Preprocess with Character Mapping: Add a step to replace common OCR-confused characters before spell-checking. For example:
This filters out obvious misrecognitions early, so your spell-checker only deals with actual spelling issues.ocr_char_mapping = {'0': 'o', '1': 'l', '5': 's', '2': 'z', '!': 'l'} def fix_ocr_chars(text): return ''.join([ocr_char_mapping.get(c, c) for c in text.lower()]) - Use a Domain-Specific Dictionary: Generic 100k word lists include irrelevant terms that slow down candidate generation. Build a custom dictionary by:
- Extracting high-frequency correct words from your valid OCR text (use
Counterto count occurrences, then filter out obvious errors). - Adding domain-specific jargon if your text is niche (e.g., medical, legal, technical).
Store this dictionary as a Pythonsetfor O(1) lookups:
# Preload once at program start custom_dict = set(line.strip().lower() for line in open('domain_dictionary.txt')) - Extracting high-frequency correct words from your valid OCR text (use
- Weighted Edit Distance: Peter Norvig's default edit distance treats all character changes equally, but OCR errors have higher probabilities for certain swaps (e.g.,
i↔l). Modify the candidate generation to prioritize these high-probability swaps first, reducing unnecessary computations.
2. Boost Algorithm & Batch Processing Efficiency
- Preload Resources: Never reload dictionaries or models for every file—load them once at the start of your script. This eliminates redundant IO and initialization overhead.
- Parallelize File Processing: Batch file tasks are perfect for parallelization. Use
concurrent.futuresto process multiple files at once with your CPU's cores:from concurrent.futures import ProcessPoolExecutor def process_single_file(file_path): with open(file_path, 'r') as f: text = f.read() # Apply OCR char fix, spell-check, etc. corrected_text = your_spell_check_function(text) output_file = os.path.join(save_path, os.path.basename(file_path)) with open(output_file, 'w') as f: f.write(corrected_text) return f"Processed {file_path}" # Run in parallel if __name__ == "__main__": file_list = [os.path.join(path, f) for f in os.listdir(path) if f.endswith('.txt')] with ProcessPoolExecutor() as executor: results = executor.map(process_single_file, file_list) for res in results: print(res) - Minimize Redundant Checks: Use
enchant(or your custom dict) to quickly verify if a word is valid before generating correction candidates. Only run the expensive candidate generation for invalid words:import enchant # Preload enchant dict once eng_dict = enchant.Dict("en_US") def check_and_correct(word): if eng_dict.check(word): return word # Generate correction candidates only if needed candidates = your_candidate_generation_function(word) return max(candidates, key=lambda w: eng_dict.check(w) or your_probability_score(w)) - Cache Correction Results: Use
functools.lru_cacheto cache results for repeated misspelled words (common in OCR text):from functools import lru_cache @lru_cache(maxsize=10000) def cached_correct(word): return check_and_correct(word)
3. Simplify Dependency Overhead
TextBlob is convenient but adds unnecessary layers if you're already using enchant or a custom dict. Stick to lighter-weight tools for your core logic:
- Replace TextBlob with direct
enchantcalls or your custom dictionary lookups to reduce overhead. - Avoid regex overhead where possible—use string operations for simple text manipulations (like the character mapping above) instead of complex regex patterns.
Here's a revised version of your initial code incorporating some of these tweaks:
import os import re import enchant from collections import Counter from concurrent.futures import ProcessPoolExecutor from functools import lru_cache # Preload resources once path = "/home/avics/PycharmProjects/spell_checker/textfile/" save_path = "/home/avics/PycharmProjects/spell_checker/output_file/" os.makedirs(save_path, exist_ok=True) eng_dict = enchant.Dict("en_US") # Optional: Load custom domain dictionary custom_dict = set() try: with open("domain_dictionary.txt", "r") as f: custom_dict = set(line.strip().lower() for line in f) except FileNotFoundError: print("Custom dictionary not found, using default enchant dict.") OCR_CHAR_MAPPING = {'0': 'o', '1': 'l', '5': 's', '2': 'z', '!': 'l', '8': 'b'} def fix_ocr_chars(text): return ''.join([OCR_CHAR_MAPPING.get(c, c) for c in text.lower()]) @lru_cache(maxsize=10000) def correct_word(word): cleaned_word = fix_ocr_chars(word) if cleaned_word in custom_dict or eng_dict.check(cleaned_word): return cleaned_word # Get enchant's suggested corrections (faster than custom candidate gen for starters) suggestions = eng_dict.suggest(cleaned_word) return suggestions[0] if suggestions else cleaned_word def spell_check_text(text): # Split text into words (adjust regex for your OCR text's formatting) words = re.findall(r'\b\w+\b', text) corrected_words = [correct_word(word) for word in words] # Reconstruct text (you might need to preserve punctuation/whitespace better) corrected_text = ' '.join(corrected_words) return corrected_text def process_file(file_path): try: with open(file_path, 'r', encoding='utf-8') as f: raw_text = f.read() corrected_text = spell_check_text(raw_text) output_filename = os.path.basename(file_path).replace('.txt', '_corrected.txt') output_path = os.path.join(save_path, output_filename) with open(output_path, 'w', encoding='utf-8') as f: f.write(corrected_text) return f"Successfully processed: {file_path}" except Exception as e: return f"Failed to process {file_path}: {str(e)}" if __name__ == "__main__": all_files = [os.path.join(path, fname) for fname in os.listdir(path) if fname.endswith('.txt')] with ProcessPoolExecutor() as executor: results = executor.map(process_file, all_files) for result in results: print(result)
内容的提问来源于stack exchange,提问作者AVICS55

