You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:优化OpenCV扫描结构化表单矩形检测,解决漏检问题

问题描述

作为OCR新手,我尝试用OpenCV检测扫描的结构化表单中的所有矩形框,但运行代码后大量矩形未被识别(未标注绿色边框和左上角编号的矩形已用红星标记)。我试过调整cv2.adaptiveThreshold的类型、块大小和常数参数,这些遗漏的矩形还是检测不到。请问我忽略了什么?该如何优化以确保所有矩形都被检测到?求优化自适应阈值及检测逻辑的方案。

原代码如下:

import cv2
import imutils
import warnings
import numpy as np

warnings.filterwarnings('ignore')
import matplotlib.pyplot as plt

img = cv2.imread("example.jpg") 
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

threshold = cv2.adaptiveThreshold(
    gray.copy(), 
    255, # maximum value assigned to pixel values exceeding the threshold
    cv2.ADAPTIVE_THRESH_GAUSSIAN_C,  # gaussian weighted sum of neighborhood
    cv2.THRESH_BINARY_INV,  # thresholding type 
    301, # block size (5x5 window)
    21) # constant

font = cv2.FONT_HERSHEY_COMPLEX
keypoints = cv2.findContours(threshold.copy(), 
                             cv2.RETR_CCOMP, 
                             cv2.CHAIN_APPROX_SIMPLE)
contours = imutils.grab_contours(keypoints)
working_image = None
idx = 1
cropped_field_images = []

contour_list = list(contours)
contour_list.reverse()
rev_contours = tuple(contour_list)

for contour in rev_contours:   
    x,y,w,h = cv2.boundingRect(contour) 
    area = cv2.contourArea(contour)
    approx = cv2.approxPolyDP(contour, 10, True)
    location = None
    if len(approx) == 4 and area > 1500 : #if the shape size is rectangular
        working_image = cv2.rectangle(img,(x,y),(x+w,y+h),(0,255,0),2)   
        cv2.putText(img, str(idx), (x, y), font, 1, (0,0,255))
        
        location = approx
        mask = np.zeros(gray.shape, np.uint8) #Create a blank mask
        rect_img = cv2.drawContours(mask, [location], 0, 255, -1) 
        rect_img = cv2.bitwise_and(img, img, mask = mask) 
        
        (x, y) = np.where(mask==255)
        (x1, y1) = (np.min(x), np.min(y))
        (x2, y2) = (np.max(x), np.max(y))
        cropped_rect = gray[x1:x2+1, y1:y2+1]
        
        cropped_field_images.append(cropped_rect)
        
        idx += 1
    
plt.figure(figsize = (11.69*2,8.27*2))
plt.axis('off')
plt.imshow(cv2.cvtColor(working_image, cv2.COLOR_BGR2RGB));

优化方案

1. 增加预处理步骤(去噪+形态学强化)

扫描表单存在的纸张纹理、扫描噪点会破坏细线矩形的边缘,直接阈值化会导致轮廓断裂。先做两步预处理:

  • 高斯模糊:用cv2.GaussianBlur平滑噪点,同时保留边框细节
  • 形态学膨胀:用矩形核膨胀图像,填补边框的微小断裂,让轮廓更连续

2. 调整自适应阈值参数

原代码的块尺寸(301)过大、常数(21)过高,导致小矩形的边框被当成背景过滤。调整为:

  • 块大小改为31(必须是奇数,适配小矩形的局部区域)
  • 常数改为3(降低阈值偏移,让细线边框更容易被保留)
  • 保留ADAPTIVE_THRESH_GAUSSIAN_C和THRESH_BINARY_INV的组合,适配表单的明暗不均问题

3. 改进轮廓检测模式

原代码用cv2.RETR_CCOMP可能遗漏嵌套的小轮廓,换成cv2.RETR_TREE保留所有层级轮廓,避免小矩形被大轮廓覆盖。

4. 优化矩形过滤条件

  • 原代码的area > 1500阈值过高,很多小矩形面积不达标,调整为area > 200(可根据表单实际尺寸微调)
  • 近似多边形的精度参数10过大,改为2,更精准识别矩形的四个顶点
  • 增加宽高比过滤:0.2 < w/h < 5,排除过窄或过高的非目标轮廓,减少误检

优化后代码
import cv2
import imutils
import warnings
import numpy as np

warnings.filterwarnings('ignore')
import matplotlib.pyplot as plt

img = cv2.imread("example.jpg") 
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

# 预处理:去噪+形态学强化
blurred = cv2.GaussianBlur(gray, (5, 5), 0)
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3, 3))
dilated = cv2.dilate(blurred, kernel, iterations=1)

# 调整自适应阈值参数
threshold = cv2.adaptiveThreshold(
    dilated.copy(), 
    255,
    cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
    cv2.THRESH_BINARY_INV,
    31,  # 缩小块尺寸
    3)   # 降低常数

font = cv2.FONT_HERSHEY_COMPLEX
# 改用RETR_TREE保留所有层级轮廓
keypoints = cv2.findContours(threshold.copy(), 
                             cv2.RETR_TREE, 
                             cv2.CHAIN_APPROX_SIMPLE)
contours = imutils.grab_contours(keypoints)
working_image = img.copy()  # 避免修改原图
idx = 1
cropped_field_images = []

# 按默认轮廓顺序处理即可
for contour in contours:   
    x,y,w,h = cv2.boundingRect(contour) 
    area = cv2.contourArea(contour)
    # 优化近似精度和过滤条件
    approx = cv2.approxPolyDP(contour, 2, True)
    if len(approx) == 4 and area > 200 and 0.2 < (w/h) < 5:
        working_image = cv2.rectangle(working_image,(x,y),(x+w,y+h),(0,255,0),2)   
        cv2.putText(working_image, str(idx), (x, y), font, 0.5, (0,0,255))  # 缩小字体避免遮挡
        
        location = approx
        mask = np.zeros(gray.shape, np.uint8)
        rect_img = cv2.drawContours(mask, [location], 0, 255, -1) 
        rect_img = cv2.bitwise_and(img, img, mask = mask) 
        
        (x_coords, y_coords) = np.where(mask==255)
        (x1, y1) = (np.min(x_coords), np.min(y_coords))
        (x2, y2) = (np.max(x_coords), np.max(y_coords))
        cropped_rect = gray[x1:x2+1, y1:y2+1]
        
        cropped_field_images.append(cropped_rect)
        
        idx += 1
    
plt.figure(figsize = (11.69*2,8.27*2))
plt.axis('off')
plt.imshow(cv2.cvtColor(working_image, cv2.COLOR_BGR2RGB));

关键改进说明

  • 预处理的高斯模糊和膨胀操作解决了扫描噪点导致的边框断裂问题,让小矩形的轮廓更连续
  • 阈值参数的调整让细线边框能被正确二值化,不会被当成背景过滤
  • 轮廓检测模式和过滤条件的优化,避免了遗漏小尺寸矩形,同时减少误检

内容的提问来源于stack exchange,提问作者Timothy Tuti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 12:30:01