You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多语言OCR提取难题:如何仅识别英文并忽略其他语言?

解决Py-Tesseract仅识别英文字符的问题

你遇到的核心问题是:即使指定lang='eng',Tesseract仍会将印地语字符识别为相似的英文字符,事后过滤无法逆转这种错误识别。要从根源解决,需要在Tesseract的识别阶段就限制仅处理英文字符集,结合优化预处理和识别参数。

关键优化方案

1. 使用Tesseract原生字符白名单参数

直接通过config参数传入-c tessedit_char_whitelist,让Tesseract在识别时只考虑白名单内的字符,彻底跳过非目标字符的识别,避免乱码生成。

2. 优化页面分割模式(PSM)

根据图片布局选择合适的--psm参数,比如--psm 6(适用于文本是单一均匀块的场景),帮助Tesseract更准确地解析文本结构。

3. 增强预处理(二值化)

对灰度图进行阈值处理,生成黑白对比强烈的图像,减少背景干扰,提升Tesseract的识别精度。

修改后的完整代码

import pytesseract
from PIL import Image
import numpy as np

def preprocess_image(image_path):
    # 打开图片并转灰度
    image = Image.open(image_path).convert("L")
    # 二值化处理(可根据实际图片调整阈值)
    img_array = np.array(image)
    threshold = 127
    binary_image = Image.fromarray(np.where(img_array > threshold, 255, 0).astype(np.uint8))
    return binary_image

def ocr_with_whitelist(image_path, whitelist):
    # 配置Tesseract参数:仅英文模型+字符白名单+页面分割模式
    custom_config = f'--psm 6 -l eng -c tessedit_char_whitelist={whitelist}'
    # 预处理图片
    preprocessed_image = preprocess_image(image_path)
    # 执行OCR
    extracted_text = pytesseract.image_to_string(preprocessed_image, config=custom_config)
    return extracted_text

image_path = 'aadhaar_front.png'
# 白名单包含所需英文字符、数字和标点
whitelist = 'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789/:=, \n'

extracted_text = ocr_with_whitelist(image_path, whitelist)
print(extracted_text)

说明

  • tessedit_char_whitelist参数直接限制Tesseract的识别字符集,从根源避免非英文字符被错误识别为乱码英文。
  • 二值化预处理能大幅提升文本与背景的对比度,减少印地语字符的干扰,让Tesseract更聚焦于目标英文字符。
  • 若图片是多段落或分散布局,可尝试--psm 3(默认模式)或其他PSM选项(参考Tesseract官方文档的PSM参数说明)。

内容的提问来源于stack exchange,提问作者Staron D'Silva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 02:42:48