You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tesseract OCR识别图像文本出错:I误判为1、9误判为3求解决方案

解决Tesseract字符识别错误(I识别为1、9识别为3)的方案

问题说明

提取图像中白色粗体文本时,Tesseract出现字符混淆:字母I被识别为数字1,数字9被识别为数字3,预期正确结果为I6M-9U。已尝试裁剪放大图像、二值化去除背景伪影,但识别错误仍存在。

初始代码:

def get_text_from_image(image: cv2.Mat) -> str:
    pytesseract.pytesseract.tesseract_cmd = r'C:\Tesseract-OCR\tesseract.exe'
    
    # Crop image to only get the piece I am interested in
    top, left, height, width = 25, 170, 40, 250

    try:
        crop_img = image[top:top + height, left:left + width]
        
        # Make it bigger
        resize_scaling = 1500
        resize_width = int(crop_img.shape[1] * resize_scaling / 100)
        resize_height = int(crop_img.shape[0] * resize_scaling / 100)
        resized_dimensions = (resize_width, resize_height)
    
        # Resize it
        crop_img = cv2.resize(crop_img, resized_dimensions, interpolation=cv2.INTER_CUBIC)
        
        return str(pytesseract.image_to_string(crop_img, config="--psm 6"))

更新后的二值化代码片段:

ret, thresh1 = cv.threshold(image, 120, 255, cv.THRESH_BINARY +
                                            cv.THRESH_OTSU)

cv.imshow("image", thresh1)

优化方案

1. 限定字符识别白名单

通过Tesseract配置参数,只允许目标字符范围(大写字母、数字、连字符),减少识别歧义:

# 在image_to_string的config中添加白名单
custom_config = r"--psm 6 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-"
return pytesseract.image_to_string(crop_img, config=custom_config)

2. 反转图像适配Tesseract默认识别逻辑

Tesseract默认更擅长识别深色文本在浅色背景,若你的目标是白色文本+深色背景,可反转图像颜色:

# 反转图像(白文本变黑,背景变白)
inverted_img = cv2.bitwise_not(crop_img)
# 结合二值化处理
ret, thresh_img = cv2.threshold(inverted_img, 120, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
# 用处理后的图像执行识别
return pytesseract.image_to_string(thresh_img, config=custom_config)

3. 调整PSM识别模式

--psm 6假设图像是单一均匀文本块,若字符特征特殊,可尝试更精准的模式:

  • --psm 8:假设图像是单个单词
  • --psm 10:假设图像是单个字符(需配合字符分割逻辑)

示例:

custom_config = r"--psm 8 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-"
return pytesseract.image_to_string(thresh_img, config=custom_config)

4. 自定义字体训练(进阶)

若以上方法均无效,可针对目标字体训练Tesseract自定义语言包,让模型熟悉该字体中I和9的独特特征。

整合后的完整代码

import cv2
import pytesseract

def get_text_from_image(image: cv2.Mat) -> str:
    pytesseract.pytesseract.tesseract_cmd = r'C:\Tesseract-OCR\tesseract.exe'
    
    # 裁剪目标区域
    top, left, height, width = 25, 170, 40, 250
    crop_img = image[top:top + height, left:left + width]
    
    # 放大图像
    resize_scaling = 1500
    resize_width = int(crop_img.shape[1] * resize_scaling / 100)
    resize_height = int(crop_img.shape[0] * resize_scaling / 100)
    resized_img = cv2.resize(crop_img, (resize_width, resize_height), interpolation=cv2.INTER_CUBIC)
    
    # 反转图像+二值化处理
    inverted_img = cv2.bitwise_not(resized_img)
    ret, thresh_img = cv2.threshold(inverted_img, 120, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
    
    # 配置白名单和PSM模式识别
    custom_config = r"--psm 6 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-"
    return pytesseract.image_to_string(thresh_img, config=custom_config).strip()

内容的提问来源于stack exchange,提问作者Andreas Ellsen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 14:03:18