You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Tesseract识别图像中的数字及其位置时出现误识别的问题求助

Tesseract识别图像中的数字及其位置时出现误识别的问题求助

我现在正尝试识别图像里的数字以及它们的位置,但遇到了不少误识别的情况——有些数字被识别成了@、|这类符号,调整psm参数也没解决问题。下面是我的代码和输出结果,想问问大家我是不是漏掉了什么关键步骤?

我的代码

import cv2
import pytesseract

def round_to_nearest_10(number):
    return round(number / 10) * 10

def parse_image_grid(filename):
    # Set the path to the Tesseract executable (update with your path)
    pytesseract.pytesseract.tesseract_cmd = r'C:\\Program Files\\Tesseract-OCR\\tesseract.exe'

    # Read the image
    image = cv2.imread(filename)

    # Convert the image to grayscale
    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

    # Apply GaussianBlur to reduce noise and improve OCR accuracy
    blurred = cv2.GaussianBlur(gray, (5, 5), 0)

    # Use the Canny edge detector to find edges in the image
    edges = cv2.Canny(blurred, 50, 150)

    # Find contours in the image
    contours, _ = cv2.findContours(edges.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)

    # Dictionary to store the mapping of square coordinates to identified numbers
    square_dict = {}

    # Iterate through each contour
    for contour in contours:
        # Approximate the contour to a polygon
        epsilon = 0.04 * cv2.arcLength(contour, True)
        approx = cv2.approxPolyDP(contour, epsilon, True)

        # Check if the polygon has four corners (likely a square)
        if len(approx) == 4:
            # Extract the region of interest (ROI) containing the square
            x, y, w, h = cv2.boundingRect(contour)
            square_roi = image[y:y + h, x:x + w]

            # print(square_roi)
            # Use OCR to extract numbers from the square
            square_text = pytesseract.image_to_string(square_roi, config="--psm 6").strip()

            # Print the square coordinates and extracted numbers
            print(f"Square at ({x}, {y}), Numbers: {square_text}")

当前输出

Square at (221, 71), Numbers: 4a
Square at (181, 61), Numbers: fi
Square at (31, 61), Numbers: 3 |
Square at (211, 31), Numbers: @
Square at (181, 31), Numbers: 2
Square at (121, 31), Numbers: ff
Square at (91, 31), Numbers: &
Square at (61, 31), Numbers: @
Square at (1, 31), Numbers:
Square at (121, 1), Numbers: 5 |
Square at (91, 1), Numbers: Es
Square at (61, 1), Numbers: @
Square at (31, 0), Numbers: 9

可以看到部分方块识别正确,但其他的数字被识别成了@、|这类无关字符,调整psm参数后也没有改善,想请教下我是不是遗漏了什么优化步骤?


可能的优化方案

我根据经验整理了几个可以尝试的方向,你可以逐个测试:

  • 强化图像预处理:目前的高斯模糊和边缘检测可能不够,建议增加二值化处理,比如用cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)生成高对比度的黑白图像,让数字轮廓更清晰;还可以尝试形态学操作(比如cv2.morphologyEx)消除小噪点,避免干扰OCR识别。
  • 限定识别字符范围:既然只需要识别数字,在Tesseract的config里加上-c tessedit_char_whitelist=0123456789,这样Tesseract只会尝试匹配数字,直接过滤掉其他字符,能大幅降低误识别概率。
  • 优化轮廓筛选逻辑:现在只判断轮廓是四边形,但很多无效区域也会被选中。可以增加面积筛选,比如计算轮廓面积cv2.contourArea(contour),设置合理的最小/最大面积阈值,过滤掉太小(噪点)或太大(非目标方块)的轮廓。
  • 尝试更适配的psm模式:--psm 6是假设图像是一个单一的均匀文本块,但你的每个ROI是单个数字,试试--psm 10(单字符识别模式)或者--psm 8(假设图像是单个词),可能更符合你的场景。
  • 检查ROI裁剪精度:有时候裁剪的ROI可能包含多余的边缘,或者没完全框住数字。可以尝试微调裁剪坐标,比如x+5, y+5, w-10, h-10(根据实际情况调整),去掉边缘的干扰区域,让数字在ROI中更突出。

备注:内容来源于stack exchange,提问作者123456789

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.20 09:44:35