使用Pytesseract无法识别燃气表数字?求预处理优化方案
燃气表数字识别预处理优化求助
我正尝试用Python和OpenCV读取燃气表数字,已定位到数字区域:
随后提取出的黑白数字图像如下:





但Pytesseract始终无法识别这些数字,试过其他分割模式也无效,推测问题出在预处理环节,求相关优化建议。
我的代码:
img = cv2.imread("gasmeter.jpg") # Resize image scale_percent = 40 width = int(img.shape[1] * scale_percent / 100) height = int(img.shape[0] * scale_percent / 100) dim = (width, height) resized_img = cv2.resize(img, dim, interpolation=cv2.INTER_AREA) grayscale = cv_funcs.get_grayscale(resized_img) thresh = cv_funcs.thresholding(grayscale) contours, hierarchy = cv2.findContours(thresh, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE) sorted_contours = [] for cnt in contours: area = cv2.contourArea(cnt) approx = cv2.approxPolyDP(cnt,0.01*cv2.arcLength(cnt,True),True) # Find number of edges of the contours if area > 840.0 and area < 1000.0 and len(approx) > 4 and len(approx) < 12: sorted_contours.append(cnt) if area > 6000.0 and area < 6500.0 and len(approx) < 12: sorted_contours.append(cnt) all_contours_img = resized_img.copy() correct_contours_img = resized_img.copy() con = cv2.drawContours(all_contours_img, contours, -1, (0,255,0), 3) con2 = cv2.drawContours(correct_contours_img, sorted_contours, -1, (0,255,0), 3) print("Finding numbers:") bounding_rect_img = resized_img.copy() # Find bounding rectangle for i,cnt in enumerate(sorted_contours): x,y,w,h = cv2.boundingRect(cnt) cv2.rectangle(bounding_rect_img,(x,y),(x+w,y+h),(0,255,0),2) # Find number crop_img = thresh[y:y+h, x:x+w] invert_img = cv2.bitwise_not(crop_img) erode_img = cv_funcs.erode(invert_img) dialate_img = cv_funcs.dilate(erode_img) cv2.imshow(str(i), erode_img) cv2.imwrite(str(i)+'.jpg', erode_img) text = pytesseract.image_to_string(erode_img, config='--psm 10 digits') print(text)
预处理优化建议
- 调整图像缩放比例:当前将图像缩小到40%会丢失数字细节,建议先放大图像(比如200%)再做后续处理,保留更多数字边缘特征,避免Tesseract无法捕捉关键识别点。
- 替换自适应阈值二值化:如果
cv_funcs.thresholding使用的是固定阈值,换成cv2.adaptiveThreshold,能更好应对局部光照不均问题,让数字与背景分离更彻底。 - 精细化形态学操作:当前的腐蚀、膨胀操作参数可能不合适,手动调整核的大小(比如用
cv2.getStructuringElement(cv2.MORPH_RECT, (2,2))),避免数字变形或残留噪点干扰识别。 - 统一数字区域尺寸:将每个裁剪后的数字图像调整到统一大小(比如28x28或更高分辨率),Tesseract对标准尺寸的字符识别准确率更高。
- 优化Tesseract配置:除
--psm 10外,添加--oem 3启用默认OCR引擎模式;尝试--psm 8(单字符识别模式);如果燃气表字体特殊,可训练自定义数字字体库匹配识别。 - 去除冗余边框干扰:从提取的数字图像看,部分数字周围有多余边框,可通过二次轮廓分析只保留数字本身的轮廓,重新生成干净的数字图像,减少无关元素的干扰。
内容的提问来源于stack exchange,提问作者tobiasrj
相关产品推荐
相关产品推荐

