You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pytesseract识别4x4数字矩阵图像失败,求解决方案

4x4数字矩阵OCR识别错误的优化方案

一、优化图像预处理,适配光照差异

原固定阈值二值化无法适配图像局部明暗差异,容易导致数字与背景分离不彻底,换成自适应阈值+高斯模糊降噪能显著提升预处理效果:

def preprocess_image(image_path):
    image = cv2.imread(image_path)
    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
    # 高斯模糊降噪,消除图像杂点干扰
    blurred = cv2.GaussianBlur(gray, (3, 3), 0)
    # 自适应二值化,根据局部区域明暗自动调整阈值
    binary = cv2.adaptiveThreshold(blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, 
                                  cv2.THRESH_BINARY_INV, 11, 2)
    return binary

二、精准筛选单元格轮廓,过滤干扰

原轮廓检测可能误抓数字小轮廓或噪点,需通过面积过滤保留单元格级别的轮廓,并确保排序后符合4x4矩阵结构:

def find_cells(binary_image):
    contours, _ = cv2.findContours(binary_image, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    # 过滤面积过小的干扰轮廓(阈值需根据实际图像尺寸调整)
    min_area = 500
    valid_contours = [cnt for cnt in contours if cv2.contourArea(cnt) > min_area]
    bounding_boxes = [cv2.boundingRect(cnt) for cnt in valid_contours]
    # 先按行分组排序,再按列排序,确保输出4行每行4个单元格
    bounding_boxes = sorted(bounding_boxes, key=lambda x: (x[1] // 100, x[0]))
    return bounding_boxes

三、优化Tesseract OCR配置,强制识别数字

原配置未限制字符范围,导致识别出字母,需指定仅识别数字,并针对单数字场景调整页面分割模式:

def extract_text_from_cells(binary_image, bounding_boxes):
    cell_texts = []
    for box in bounding_boxes:
        left, top, width, height = box
        # 裁剪时去掉边框干扰,保留核心数字区域
        margin = 5
        cropped = binary_image[top+margin:top+height-margin, left+margin:left+width-margin]
        # OCR配置:仅识别数字,单字符识别模式
        text = pytesseract.image_to_string(cropped, config='--oem 3 --psm 10 -c tessedit_char_whitelist=0123456789')
        cell_texts.append(text.strip())
    return cell_texts

完整优化后代码

import cv2
import numpy as np
import pytesseract

def preprocess_image(image_path):
    image = cv2.imread(image_path)
    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
    blurred = cv2.GaussianBlur(gray, (3, 3), 0)
    binary = cv2.adaptiveThreshold(blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, 
                                  cv2.THRESH_BINARY_INV, 11, 2)
    return binary

def find_cells(binary_image):
    contours, _ = cv2.findContours(binary_image, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    min_area = 500
    valid_contours = [cnt for cnt in contours if cv2.contourArea(cnt) > min_area]
    bounding_boxes = [cv2.boundingRect(cnt) for cnt in valid_contours]
    bounding_boxes = sorted(bounding_boxes, key=lambda x: (x[1] // 100, x[0]))
    return bounding_boxes

def extract_text_from_cells(binary_image, bounding_boxes):
    cell_texts = []
    for box in bounding_boxes:
        left, top, width, height = box
        margin = 5
        cropped = binary_image[top+margin:top+height-margin, left+margin:left+width-margin]
        text = pytesseract.image_to_string(cropped, config='--oem 3 --psm 10 -c tessedit_char_whitelist=0123456789')
        cell_texts.append(text.strip())
    return cell_texts

image_path = r'C:\Users\Sandro\Desktop\BINGO\screenshot.png'

binary_image = preprocess_image(image_path)
cells = find_cells(binary_image)
texts = extract_text_from_cells(binary_image, cells)

# 按4x4矩阵格式输出结果
for i in range(0, len(texts), 4):
    print(f"第{i//4 +1}行: {texts[i:i+4]}")

内容的提问来源于stack exchange,提问作者Sandro Pinho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 05:49:52