You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用OpenCV+PyTesseract准确识别图片底部的白色数字

优化数字识别的解决方案

一、优化OpenCV预处理步骤

针对数字识别的常见问题,调整预处理流程增强数字特征:

  • 降噪处理:先对灰度图做高斯模糊,消除微小噪点干扰
  • 自适应二值化:替代固定阈值二值化,适配局部光照差异
  • 形态学操作:用闭运算填充数字内部空隙、连接断裂笔画

优化后的预处理代码:

import cv2
import numpy as np
import math
import pytesseract
import logging

img = cv2.imread(local_file_path)
gray_image = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

# 高斯模糊降噪
blurred = cv2.GaussianBlur(gray_image, (3, 3), 0)

# 自适应二值化,适配局部光照
thresh = cv2.adaptiveThreshold(blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, 
                               cv2.THRESH_BINARY_INV, 11, 2)

# 形态学闭运算,修复数字笔画断裂
kernel = np.ones((2, 2), np.uint8)
thresh = cv2.morphologyEx(thresh, cv2.MORPH_CLOSE, kernel, iterations=1)

# 后续ROI裁剪逻辑不变
h, w, c = img.shape
columns = math.ceil(w / 255)

for column_num in range(columns):
    roi_x_1 = 45 + (column_num * 255)
    roi_y_1 = h - 35
    roi_x_2 = 80 + (column_num * 255)
    roi_y_2 = h
    roi = thresh[roi_y_1:roi_y_2, roi_x_1:roi_x_2]
    
    # 可选:放大ROI,提升小数字识别率
    roi = cv2.resize(roi, None, fx=2, fy=2, interpolation=cv2.INTER_CUBIC)
    
    # 调整Tesseract配置
    config = '--psm 7 --oem 1 -c tessedit_char_whitelist=0123456789'
    data = pytesseract.image_to_string(roi, config=config).strip()
    
    if data.isdigit():
        measure_number = int(data)
        logging.warning(f"Number found in column {column_num + 1}: {measure_number}")
    else:
        logging.warning(f"Failed to recognize number in column {column_num + 1}")

二、调整Tesseract配置

  • 引擎模式(OEM):用--oem 1指定仅使用LSTM引擎,Tesseract 4+的LSTM对数字识别准确率远高于传统引擎
  • 页面分割模式(PSM):用--psm 7(单文本行模式)或--psm 8(单单词模式),比--psm 6更适配短数字场景
  • 结果清洗:识别后用strip()去除空白字符,再判断是否为纯数字,避免转换报错

三、备选方案:使用EasyOCR

如果Tesseract仍无法满足需求,EasyOCR对数字的原生识别效果更优,且无需复杂预处理:

import easyocr
import cv2
import math
import logging

reader = easyocr.Reader(['en'], gpu=False)  # 有GPU可设为True加速

img = cv2.imread(local_file_path)
h, w, c = img.shape
columns = math.ceil(w / 255)

for column_num in range(columns):
    roi_x_1 = 45 + (column_num * 255)
    roi_y_1 = h - 35
    roi_x_2 = 80 + (column_num * 255)
    roi_y_2 = h
    roi = img[roi_y_1:roi_y_2, roi_x_1:roi_x_2]
    
    # 仅识别数字
    result = reader.readtext(roi, allowlist='0123456789')
    if result:
        measure_number = int(result[0][1])
        logging.warning(f"Number found in column {column_num + 1}: {measure_number}")
    else:
        logging.warning(f"Failed to recognize number in column {column_num + 1}")

四、辅助排查建议

  • 循环中添加cv2.imwrite(f'roi_{column_num}.png', roi),手动检查每个ROI的数字是否清晰、完整
  • 若数字存在倾斜,可添加旋转矫正步骤(如霍夫变换检测直线后旋转)

内容的提问来源于stack exchange,提问作者Axel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 18:27:34