Pytesseract无法识别小尺寸裁剪图片中的数字,求解决方案
问题描述
有一张从原图裁剪得到的小尺寸数字图片,尝试图片放大、阈值处理等预处理操作后,仍无法通过Pytesseract识别其中的数字。
原识别代码:
import cv2 import pytesseract from pytesseract import Output img = cv2.imread('rois/roi11.jpg') data = pytesseract.image_to_boxes(img, output_type=Output.DICT) print(data)
尝试的预处理代码:
import cv2 import pytesseract img = cv2.imread('rois/roi11.jpg') img2 = cv2.resize(img, (0, 0), fx=2, fy=2) gry = cv2.cvtColor(img2, cv2.COLOR_BGR2GRAY) thr = cv2.threshold(gry, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)[1] data = pytesseract.image_to_string(thr) print(data)
可行解决方法
- 指定数字识别模式:Pytesseract默认识别全字符,通过
--psm 10(单字符识别模式)和-c tessedit_char_whitelist=0123456789(限定仅识别数字)缩小识别范围,提升精准度。同时改用INTER_CUBIC插值放大,保留更多字符细节:
import cv2 import pytesseract img = cv2.imread('rois/roi11.jpg') # 用三次插值放大4倍,比线性插值更清晰 img_enlarged = cv2.resize(img, None, fx=4, fy=4, interpolation=cv2.INTER_CUBIC) gray = cv2.cvtColor(img_enlarged, cv2.COLOR_BGR2GRAY) # 自适应阈值处理,适配局部明暗差异 thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) # 设置自定义识别参数 custom_config = r'--oem 3 --psm 10 -c tessedit_char_whitelist=0123456789' result = pytesseract.image_to_string(thresh, config=custom_config) print(result.strip())
- 形态学优化轮廓:对阈值处理后的图像进行膨胀操作,填补数字边缘的细小缺口,强化字符轮廓:
import cv2 import pytesseract import numpy as np img = cv2.imread('rois/roi11.jpg') img_enlarged = cv2.resize(img, None, fx=4, fy=4, interpolation=cv2.INTER_CUBIC) gray = cv2.cvtColor(img_enlarged, cv2.COLOR_BGR2GRAY) thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) # 创建2x2结构元素,执行膨胀操作 kernel = np.ones((2, 2), np.uint8) dilated_img = cv2.dilate(thresh, kernel, iterations=1) custom_config = r'--oem 3 --psm 10 -c tessedit_char_whitelist=0123456789' result = pytesseract.image_to_string(dilated_img, config=custom_config) print(result.strip())
- 更换OCR引擎模式:若你的Tesseract版本支持,设置
--oem 1启用LSTM神经网络引擎,对模糊、小字符的识别效果优于传统引擎:
custom_config = r'--oem 1 --psm 10 -c tessedit_char_whitelist=0123456789'
- 降噪预处理:如果图片存在明显噪点,先通过高斯模糊降噪,再进行后续处理:
gray = cv2.cvtColor(img_enlarged, cv2.COLOR_BGR2GRAY) # 高斯模糊降噪 blurred = cv2.GaussianBlur(gray, (3, 3), 0) thresh = cv2.adaptiveThreshold(blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
内容的提问来源于stack exchange,提问作者naiveprogrammer
相关产品推荐
相关产品推荐

