多语言OCR提取难题:如何仅识别英文并忽略其他语言?
解决Py-Tesseract仅识别英文字符的问题
你遇到的核心问题是:即使指定lang='eng',Tesseract仍会将印地语字符识别为相似的英文字符,事后过滤无法逆转这种错误识别。要从根源解决,需要在Tesseract的识别阶段就限制仅处理英文字符集,结合优化预处理和识别参数。
关键优化方案
1. 使用Tesseract原生字符白名单参数
直接通过config参数传入-c tessedit_char_whitelist,让Tesseract在识别时只考虑白名单内的字符,彻底跳过非目标字符的识别,避免乱码生成。
2. 优化页面分割模式(PSM)
根据图片布局选择合适的--psm参数,比如--psm 6(适用于文本是单一均匀块的场景),帮助Tesseract更准确地解析文本结构。
3. 增强预处理(二值化)
对灰度图进行阈值处理,生成黑白对比强烈的图像,减少背景干扰,提升Tesseract的识别精度。
修改后的完整代码
import pytesseract from PIL import Image import numpy as np def preprocess_image(image_path): # 打开图片并转灰度 image = Image.open(image_path).convert("L") # 二值化处理(可根据实际图片调整阈值) img_array = np.array(image) threshold = 127 binary_image = Image.fromarray(np.where(img_array > threshold, 255, 0).astype(np.uint8)) return binary_image def ocr_with_whitelist(image_path, whitelist): # 配置Tesseract参数:仅英文模型+字符白名单+页面分割模式 custom_config = f'--psm 6 -l eng -c tessedit_char_whitelist={whitelist}' # 预处理图片 preprocessed_image = preprocess_image(image_path) # 执行OCR extracted_text = pytesseract.image_to_string(preprocessed_image, config=custom_config) return extracted_text image_path = 'aadhaar_front.png' # 白名单包含所需英文字符、数字和标点 whitelist = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789/:=, \n' extracted_text = ocr_with_whitelist(image_path, whitelist) print(extracted_text)
说明
tessedit_char_whitelist参数直接限制Tesseract的识别字符集,从根源避免非英文字符被错误识别为乱码英文。- 二值化预处理能大幅提升文本与背景的对比度,减少印地语字符的干扰,让Tesseract更聚焦于目标英文字符。
- 若图片是多段落或分散布局,可尝试
--psm 3(默认模式)或其他PSM选项(参考Tesseract官方文档的PSM参数说明)。
内容的提问来源于stack exchange,提问作者Staron D'Silva
相关产品推荐
相关产品推荐

