使用Tesseract识别图像中的数字及其位置时出现误识别的问题求助
Tesseract识别图像中的数字及其位置时出现误识别的问题求助
我现在正尝试识别图像里的数字以及它们的位置,但遇到了不少误识别的情况——有些数字被识别成了@、|这类符号,调整psm参数也没解决问题。下面是我的代码和输出结果,想问问大家我是不是漏掉了什么关键步骤?
我的代码
import cv2 import pytesseract def round_to_nearest_10(number): return round(number / 10) * 10 def parse_image_grid(filename): # Set the path to the Tesseract executable (update with your path) pytesseract.pytesseract.tesseract_cmd = r'C:\\Program Files\\Tesseract-OCR\\tesseract.exe' # Read the image image = cv2.imread(filename) # Convert the image to grayscale gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) # Apply GaussianBlur to reduce noise and improve OCR accuracy blurred = cv2.GaussianBlur(gray, (5, 5), 0) # Use the Canny edge detector to find edges in the image edges = cv2.Canny(blurred, 50, 150) # Find contours in the image contours, _ = cv2.findContours(edges.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # Dictionary to store the mapping of square coordinates to identified numbers square_dict = {} # Iterate through each contour for contour in contours: # Approximate the contour to a polygon epsilon = 0.04 * cv2.arcLength(contour, True) approx = cv2.approxPolyDP(contour, epsilon, True) # Check if the polygon has four corners (likely a square) if len(approx) == 4: # Extract the region of interest (ROI) containing the square x, y, w, h = cv2.boundingRect(contour) square_roi = image[y:y + h, x:x + w] # print(square_roi) # Use OCR to extract numbers from the square square_text = pytesseract.image_to_string(square_roi, config="--psm 6").strip() # Print the square coordinates and extracted numbers print(f"Square at ({x}, {y}), Numbers: {square_text}")
当前输出
Square at (221, 71), Numbers: 4a Square at (181, 61), Numbers: fi Square at (31, 61), Numbers: 3 | Square at (211, 31), Numbers: @ Square at (181, 31), Numbers: 2 Square at (121, 31), Numbers: ff Square at (91, 31), Numbers: & Square at (61, 31), Numbers: @ Square at (1, 31), Numbers: Square at (121, 1), Numbers: 5 | Square at (91, 1), Numbers: Es Square at (61, 1), Numbers: @ Square at (31, 0), Numbers: 9
可以看到部分方块识别正确,但其他的数字被识别成了@、|这类无关字符,调整psm参数后也没有改善,想请教下我是不是遗漏了什么优化步骤?
可能的优化方案
我根据经验整理了几个可以尝试的方向,你可以逐个测试:
- 强化图像预处理:目前的高斯模糊和边缘检测可能不够,建议增加二值化处理,比如用
cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)生成高对比度的黑白图像,让数字轮廓更清晰;还可以尝试形态学操作(比如cv2.morphologyEx)消除小噪点,避免干扰OCR识别。 - 限定识别字符范围:既然只需要识别数字,在Tesseract的config里加上
-c tessedit_char_whitelist=0123456789,这样Tesseract只会尝试匹配数字,直接过滤掉其他字符,能大幅降低误识别概率。 - 优化轮廓筛选逻辑:现在只判断轮廓是四边形,但很多无效区域也会被选中。可以增加面积筛选,比如计算轮廓面积
cv2.contourArea(contour),设置合理的最小/最大面积阈值,过滤掉太小(噪点)或太大(非目标方块)的轮廓。 - 尝试更适配的psm模式:
--psm 6是假设图像是一个单一的均匀文本块,但你的每个ROI是单个数字,试试--psm 10(单字符识别模式)或者--psm 8(假设图像是单个词),可能更符合你的场景。 - 检查ROI裁剪精度:有时候裁剪的ROI可能包含多余的边缘,或者没完全框住数字。可以尝试微调裁剪坐标,比如
x+5, y+5, w-10, h-10(根据实际情况调整),去掉边缘的干扰区域,让数字在ROI中更突出。
备注:内容来源于stack exchange,提问作者123456789
相关产品推荐
相关产品推荐

