如何用OpenCV+PyTesseract准确识别图片底部的白色数字
优化数字识别的解决方案
一、优化OpenCV预处理步骤
针对数字识别的常见问题,调整预处理流程增强数字特征:
- 降噪处理:先对灰度图做高斯模糊,消除微小噪点干扰
- 自适应二值化:替代固定阈值二值化,适配局部光照差异
- 形态学操作:用闭运算填充数字内部空隙、连接断裂笔画
优化后的预处理代码:
import cv2 import numpy as np import math import pytesseract import logging img = cv2.imread(local_file_path) gray_image = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 高斯模糊降噪 blurred = cv2.GaussianBlur(gray_image, (3, 3), 0) # 自适应二值化,适配局部光照 thresh = cv2.adaptiveThreshold(blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) # 形态学闭运算,修复数字笔画断裂 kernel = np.ones((2, 2), np.uint8) thresh = cv2.morphologyEx(thresh, cv2.MORPH_CLOSE, kernel, iterations=1) # 后续ROI裁剪逻辑不变 h, w, c = img.shape columns = math.ceil(w / 255) for column_num in range(columns): roi_x_1 = 45 + (column_num * 255) roi_y_1 = h - 35 roi_x_2 = 80 + (column_num * 255) roi_y_2 = h roi = thresh[roi_y_1:roi_y_2, roi_x_1:roi_x_2] # 可选:放大ROI,提升小数字识别率 roi = cv2.resize(roi, None, fx=2, fy=2, interpolation=cv2.INTER_CUBIC) # 调整Tesseract配置 config = '--psm 7 --oem 1 -c tessedit_char_whitelist=0123456789' data = pytesseract.image_to_string(roi, config=config).strip() if data.isdigit(): measure_number = int(data) logging.warning(f"Number found in column {column_num + 1}: {measure_number}") else: logging.warning(f"Failed to recognize number in column {column_num + 1}")
二、调整Tesseract配置
- 引擎模式(OEM):用
--oem 1指定仅使用LSTM引擎,Tesseract 4+的LSTM对数字识别准确率远高于传统引擎 - 页面分割模式(PSM):用
--psm 7(单文本行模式)或--psm 8(单单词模式),比--psm 6更适配短数字场景 - 结果清洗:识别后用
strip()去除空白字符,再判断是否为纯数字,避免转换报错
三、备选方案:使用EasyOCR
如果Tesseract仍无法满足需求,EasyOCR对数字的原生识别效果更优,且无需复杂预处理:
import easyocr import cv2 import math import logging reader = easyocr.Reader(['en'], gpu=False) # 有GPU可设为True加速 img = cv2.imread(local_file_path) h, w, c = img.shape columns = math.ceil(w / 255) for column_num in range(columns): roi_x_1 = 45 + (column_num * 255) roi_y_1 = h - 35 roi_x_2 = 80 + (column_num * 255) roi_y_2 = h roi = img[roi_y_1:roi_y_2, roi_x_1:roi_x_2] # 仅识别数字 result = reader.readtext(roi, allowlist='0123456789') if result: measure_number = int(result[0][1]) logging.warning(f"Number found in column {column_num + 1}: {measure_number}") else: logging.warning(f"Failed to recognize number in column {column_num + 1}")
四、辅助排查建议
- 循环中添加
cv2.imwrite(f'roi_{column_num}.png', roi),手动检查每个ROI的数字是否清晰、完整 - 若数字存在倾斜,可添加旋转矫正步骤(如霍夫变换检测直线后旋转)
内容的提问来源于stack exchange,提问作者Axel
相关产品推荐
相关产品推荐

