You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于OpenCV与Pytesseract的燃气表数字OCR识别难题求助

燃气表数字识别优化方案

问题背景

需开发Python脚本提取燃气表固定位置的8位数字,图片分明亮、昏暗两种类型,每10分钟拍摄一次,拥有大量样本。当前使用Pytesseract+OpenCV2,现有方案(裁剪目标区域后调用pytesseract.image_to_string()并配置--psm 7)可靠性极低;尝试阈值处理、自适应阈值、逐字符识别均未达到预期效果,且无法优化拍摄条件。

改进建议

1. 明暗图片差异化预处理

先通过灰度图均值判断图片亮度,再针对性处理:

  • 昏暗图片:使用限制对比度自适应直方图均衡化(CLAHE)提升局部对比度,避免全局均衡化导致的过曝:
    clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8))
    gray_img = cv2.cvtColor(cropped_image, cv2.COLOR_BGR2GRAY)
    enhanced_img = clahe.apply(gray_img)
    
  • 明亮图片:采用Otsu自动阈值二值化,自动适配不同亮度:
    gray_img = cv2.cvtColor(cropped_image, cv2.COLOR_BGR2GRAY)
    _, binary_img = cv2.threshold(gray_img, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
    

2. 优化Tesseract识别配置

调整参数减少干扰,提升识别稳定性:

  • 使用--psm 6(假设输入为单一均匀文本块)替代--psm 7,更适配固定位置的数字区域
  • 指定仅识别数字的字符白名单,同时启用默认引擎模式:
    config = "--oem 3 --psm 6 -c tessedit_char_whitelist=0123456789"
    text = pytesseract.image_to_string(processed_img, config=config)
    

3. 形态学处理强化字符轮廓

对预处理后的二值图做形态学操作,消除噪声并填补字符缺口:

# 构造小型结构元素
kernel = np.ones((2, 2), np.uint8)
# 闭运算填补字符内部小缺口
closed_img = cv2.morphologyEx(binary_img, cv2.MORPH_CLOSE, kernel)
# 开运算去除背景噪声
opened_img = cv2.morphologyEx(closed_img, cv2.MORPH_OPEN, kernel)

4. 训练Tesseract自定义数字模型

利用大量样本训练专属识别模型:

  1. 从样本中裁剪出单个数字,生成.tif图片与对应.box标注文件(标注每个数字的位置与内容)
  2. 使用Tesseract训练工具(tesstrain)训练自定义数字识别模型,替换默认模型提升精度

5. 尝试替代OCR工具

若Tesseract效果仍不理想,可尝试:

  • EasyOCR:轻量易用,对小文本识别适配性好
  • PaddleOCR:开源且对数字识别精度高,支持快速部署

用户现有代码

方案一代码

import cv2
import numpy as np
import os
import pytesseract

pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract"
directory = r"C:\Users\user\Desktop\test_pcs\test"

for image in os.listdir(directory):
    
    OriginalImagePath = os.path.join(directory, image)
    OriginalImage = cv2.imread(OriginalImagePath)
    x_start, y_start = int(1110), int(445)
    x_end, y_end = int(1690), int(520)
    cropped_image = OriginalImage[y_start:y_end, x_start:x_end]
    text = (pytesseract.image_to_string(cropped_image, config="--psm 7 outputbase digits"))
    cv2.imshow("Cropped", cropped_image)
    cv2.waitKey(0)
    print(text + "    " + OriginalImagePath)
    
cv2.destroyAllWindows()

阈值处理代码

import cv2 as cv
import numpy as np
from matplotlib import pyplot as plt
import pytesseract

pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract"

img = cv.imread(r"C:\Users\user\Desktop\test_pcs\new2\2022-10-30_14-49-30.jpg",0)
img = cv.medianBlur(img,5)
ret,th1 = cv.threshold(img,127,255,cv.THRESH_BINARY)
#'Adaptive Mean Thresholding'
th2 = cv.adaptiveThreshold(img,255,cv.ADAPTIVE_THRESH_MEAN_C,\
            cv.THRESH_BINARY,11,2)
#'Adaptive Gaussian Thresholding'
th3 = cv.adaptiveThreshold(img,255,cv.ADAPTIVE_THRESH_GAUSSIAN_C,\
            cv.THRESH_BINARY,11,2)

images = [img, th2, th3]
for i in range(3):
    plt.subplot(2,2,i+1),plt.imshow(images[i],'gray')

plt.show()

x_start, y_start = int(1110), int(450)
x_end, y_end = int(1690), int(520)
cropped_image = th2[y_start:y_end, x_start:x_end]

plt.imshow(cropped_image,'gray')

text = (pytesseract.image_to_string(cropped_image, config="--psm 7 outputbase digits"))

print("digits: " + text)

内容的提问来源于stack exchange,提问作者BMolics

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 02:20:30