使用Tesseract OCR识别图像文本出错:I误判为1、9误判为3求解决方案
解决Tesseract字符识别错误(I识别为1、9识别为3)的方案
问题说明
提取图像中白色粗体文本时,Tesseract出现字符混淆:字母I被识别为数字1,数字9被识别为数字3,预期正确结果为I6M-9U。已尝试裁剪放大图像、二值化去除背景伪影,但识别错误仍存在。
初始代码:
def get_text_from_image(image: cv2.Mat) -> str: pytesseract.pytesseract.tesseract_cmd = r'C:\Tesseract-OCR\tesseract.exe' # Crop image to only get the piece I am interested in top, left, height, width = 25, 170, 40, 250 try: crop_img = image[top:top + height, left:left + width] # Make it bigger resize_scaling = 1500 resize_width = int(crop_img.shape[1] * resize_scaling / 100) resize_height = int(crop_img.shape[0] * resize_scaling / 100) resized_dimensions = (resize_width, resize_height) # Resize it crop_img = cv2.resize(crop_img, resized_dimensions, interpolation=cv2.INTER_CUBIC) return str(pytesseract.image_to_string(crop_img, config="--psm 6"))
更新后的二值化代码片段:
ret, thresh1 = cv.threshold(image, 120, 255, cv.THRESH_BINARY + cv.THRESH_OTSU) cv.imshow("image", thresh1)
优化方案
1. 限定字符识别白名单
通过Tesseract配置参数,只允许目标字符范围(大写字母、数字、连字符),减少识别歧义:
# 在image_to_string的config中添加白名单 custom_config = r"--psm 6 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-" return pytesseract.image_to_string(crop_img, config=custom_config)
2. 反转图像适配Tesseract默认识别逻辑
Tesseract默认更擅长识别深色文本在浅色背景,若你的目标是白色文本+深色背景,可反转图像颜色:
# 反转图像(白文本变黑,背景变白) inverted_img = cv2.bitwise_not(crop_img) # 结合二值化处理 ret, thresh_img = cv2.threshold(inverted_img, 120, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) # 用处理后的图像执行识别 return pytesseract.image_to_string(thresh_img, config=custom_config)
3. 调整PSM识别模式
--psm 6假设图像是单一均匀文本块,若字符特征特殊,可尝试更精准的模式:
--psm 8:假设图像是单个单词--psm 10:假设图像是单个字符(需配合字符分割逻辑)
示例:
custom_config = r"--psm 8 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-" return pytesseract.image_to_string(thresh_img, config=custom_config)
4. 自定义字体训练(进阶)
若以上方法均无效,可针对目标字体训练Tesseract自定义语言包,让模型熟悉该字体中I和9的独特特征。
整合后的完整代码
import cv2 import pytesseract def get_text_from_image(image: cv2.Mat) -> str: pytesseract.pytesseract.tesseract_cmd = r'C:\Tesseract-OCR\tesseract.exe' # 裁剪目标区域 top, left, height, width = 25, 170, 40, 250 crop_img = image[top:top + height, left:left + width] # 放大图像 resize_scaling = 1500 resize_width = int(crop_img.shape[1] * resize_scaling / 100) resize_height = int(crop_img.shape[0] * resize_scaling / 100) resized_img = cv2.resize(crop_img, (resize_width, resize_height), interpolation=cv2.INTER_CUBIC) # 反转图像+二值化处理 inverted_img = cv2.bitwise_not(resized_img) ret, thresh_img = cv2.threshold(inverted_img, 120, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) # 配置白名单和PSM模式识别 custom_config = r"--psm 6 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-" return pytesseract.image_to_string(thresh_img, config=custom_config).strip()
内容的提问来源于stack exchange,提问作者Andreas Ellsen
相关产品推荐
相关产品推荐

