You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pytesseract无法识别燃气表数字?求预处理优化方案

燃气表数字识别预处理优化求助

我正尝试用Python和OpenCV读取燃气表数字,已定位到数字区域:
燃气表数字区域

随后提取出的黑白数字图像如下:
数字图像1
数字图像2
数字图像3
数字图像4
数字图像5
数字图像6

但Pytesseract始终无法识别这些数字,试过其他分割模式也无效,推测问题出在预处理环节,求相关优化建议。

我的代码:

img = cv2.imread("gasmeter.jpg")

# Resize image
scale_percent = 40
width = int(img.shape[1] * scale_percent / 100)
height = int(img.shape[0] * scale_percent / 100)
dim = (width, height)
resized_img = cv2.resize(img, dim, interpolation=cv2.INTER_AREA)

grayscale = cv_funcs.get_grayscale(resized_img)
thresh = cv_funcs.thresholding(grayscale)

contours, hierarchy = cv2.findContours(thresh, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE)

sorted_contours = []
for cnt in contours:
    area = cv2.contourArea(cnt)
    approx = cv2.approxPolyDP(cnt,0.01*cv2.arcLength(cnt,True),True)    # Find number of edges of the contours

    if area > 840.0 and area < 1000.0 and len(approx) > 4 and len(approx) < 12:
        sorted_contours.append(cnt)

    if area > 6000.0 and area < 6500.0 and len(approx) < 12:
        sorted_contours.append(cnt)

all_contours_img = resized_img.copy()
correct_contours_img = resized_img.copy()
con = cv2.drawContours(all_contours_img, contours, -1, (0,255,0), 3)
con2 = cv2.drawContours(correct_contours_img, sorted_contours, -1, (0,255,0), 3)

print("Finding numbers:")
bounding_rect_img = resized_img.copy()
# Find bounding rectangle
for i,cnt in enumerate(sorted_contours):
    x,y,w,h = cv2.boundingRect(cnt)
    cv2.rectangle(bounding_rect_img,(x,y),(x+w,y+h),(0,255,0),2)
    # Find number
    crop_img = thresh[y:y+h, x:x+w]
    invert_img = cv2.bitwise_not(crop_img)
    erode_img = cv_funcs.erode(invert_img)
    dialate_img = cv_funcs.dilate(erode_img)
    cv2.imshow(str(i), erode_img)
    cv2.imwrite(str(i)+'.jpg', erode_img) 
    text = pytesseract.image_to_string(erode_img, config='--psm 10 digits')
    print(text)
预处理优化建议
  • 调整图像缩放比例:当前将图像缩小到40%会丢失数字细节,建议先放大图像(比如200%)再做后续处理,保留更多数字边缘特征,避免Tesseract无法捕捉关键识别点。
  • 替换自适应阈值二值化:如果cv_funcs.thresholding使用的是固定阈值,换成cv2.adaptiveThreshold,能更好应对局部光照不均问题,让数字与背景分离更彻底。
  • 精细化形态学操作:当前的腐蚀、膨胀操作参数可能不合适,手动调整核的大小(比如用cv2.getStructuringElement(cv2.MORPH_RECT, (2,2))),避免数字变形或残留噪点干扰识别。
  • 统一数字区域尺寸:将每个裁剪后的数字图像调整到统一大小(比如28x28或更高分辨率),Tesseract对标准尺寸的字符识别准确率更高。
  • 优化Tesseract配置:除--psm 10外,添加--oem 3启用默认OCR引擎模式;尝试--psm 8(单字符识别模式);如果燃气表字体特殊,可训练自定义数字字体库匹配识别。
  • 去除冗余边框干扰:从提取的数字图像看,部分数字周围有多余边框,可通过二次轮廓分析只保留数字本身的轮廓,重新生成干净的数字图像,减少无关元素的干扰。

内容的提问来源于stack exchange,提问作者tobiasrj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 09:30:51