基于OpenCV与Pytesseract的燃气表数字OCR识别难题求助
燃气表数字识别优化方案
问题背景
需开发Python脚本提取燃气表固定位置的8位数字,图片分明亮、昏暗两种类型,每10分钟拍摄一次,拥有大量样本。当前使用Pytesseract+OpenCV2,现有方案(裁剪目标区域后调用pytesseract.image_to_string()并配置--psm 7)可靠性极低;尝试阈值处理、自适应阈值、逐字符识别均未达到预期效果,且无法优化拍摄条件。
改进建议
1. 明暗图片差异化预处理
先通过灰度图均值判断图片亮度,再针对性处理:
- 昏暗图片:使用限制对比度自适应直方图均衡化(CLAHE)提升局部对比度,避免全局均衡化导致的过曝:
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8,8)) gray_img = cv2.cvtColor(cropped_image, cv2.COLOR_BGR2GRAY) enhanced_img = clahe.apply(gray_img) - 明亮图片:采用Otsu自动阈值二值化,自动适配不同亮度:
gray_img = cv2.cvtColor(cropped_image, cv2.COLOR_BGR2GRAY) _, binary_img = cv2.threshold(gray_img, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
2. 优化Tesseract识别配置
调整参数减少干扰,提升识别稳定性:
- 使用
--psm 6(假设输入为单一均匀文本块)替代--psm 7,更适配固定位置的数字区域 - 指定仅识别数字的字符白名单,同时启用默认引擎模式:
config = "--oem 3 --psm 6 -c tessedit_char_whitelist=0123456789" text = pytesseract.image_to_string(processed_img, config=config)
3. 形态学处理强化字符轮廓
对预处理后的二值图做形态学操作,消除噪声并填补字符缺口:
# 构造小型结构元素 kernel = np.ones((2, 2), np.uint8) # 闭运算填补字符内部小缺口 closed_img = cv2.morphologyEx(binary_img, cv2.MORPH_CLOSE, kernel) # 开运算去除背景噪声 opened_img = cv2.morphologyEx(closed_img, cv2.MORPH_OPEN, kernel)
4. 训练Tesseract自定义数字模型
利用大量样本训练专属识别模型:
- 从样本中裁剪出单个数字,生成
.tif图片与对应.box标注文件(标注每个数字的位置与内容) - 使用Tesseract训练工具(
tesstrain)训练自定义数字识别模型,替换默认模型提升精度
5. 尝试替代OCR工具
若Tesseract效果仍不理想,可尝试:
- EasyOCR:轻量易用,对小文本识别适配性好
- PaddleOCR:开源且对数字识别精度高,支持快速部署
用户现有代码
方案一代码
import cv2 import numpy as np import os import pytesseract pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract" directory = r"C:\Users\user\Desktop\test_pcs\test" for image in os.listdir(directory): OriginalImagePath = os.path.join(directory, image) OriginalImage = cv2.imread(OriginalImagePath) x_start, y_start = int(1110), int(445) x_end, y_end = int(1690), int(520) cropped_image = OriginalImage[y_start:y_end, x_start:x_end] text = (pytesseract.image_to_string(cropped_image, config="--psm 7 outputbase digits")) cv2.imshow("Cropped", cropped_image) cv2.waitKey(0) print(text + " " + OriginalImagePath) cv2.destroyAllWindows()
阈值处理代码
import cv2 as cv import numpy as np from matplotlib import pyplot as plt import pytesseract pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract" img = cv.imread(r"C:\Users\user\Desktop\test_pcs\new2\2022-10-30_14-49-30.jpg",0) img = cv.medianBlur(img,5) ret,th1 = cv.threshold(img,127,255,cv.THRESH_BINARY) #'Adaptive Mean Thresholding' th2 = cv.adaptiveThreshold(img,255,cv.ADAPTIVE_THRESH_MEAN_C,\ cv.THRESH_BINARY,11,2) #'Adaptive Gaussian Thresholding' th3 = cv.adaptiveThreshold(img,255,cv.ADAPTIVE_THRESH_GAUSSIAN_C,\ cv.THRESH_BINARY,11,2) images = [img, th2, th3] for i in range(3): plt.subplot(2,2,i+1),plt.imshow(images[i],'gray') plt.show() x_start, y_start = int(1110), int(450) x_end, y_end = int(1690), int(520) cropped_image = th2[y_start:y_end, x_start:x_end] plt.imshow(cropped_image,'gray') text = (pytesseract.image_to_string(cropped_image, config="--psm 7 outputbase digits")) print("digits: " + text)
内容的提问来源于stack exchange,提问作者BMolics
相关产品推荐
相关产品推荐

