Python OpenCV OMR答题卡填涂检测的二值化优化方法求解
Python+OpenCV OMR答题卡识别优化方案
问题说明
- 基于Python结合OpenCV开发OMR答题卡识别功能时,无法准确检测全部已填涂答案,使用的输入图像:输入答题卡
- 初始版本采用固定阈值做二值化处理,代码如下:
import cv2 image = cv2.imread('input.png') img = cv2.GaussianBlur(image,(5,5),0) res, img = cv2.threshold(img, 60, 255, cv2.THRESH_BINARY) img = 255 - img cv2.imwrite('output.png',img)
输出结果存在大量明显噪点:初次二值化结果
- 调整高斯模糊核与固定阈值参数后效果仍不达标:
img = cv2.GaussianBlur(image,(7,7),0) res, img = cv2.threshold(img, 90, 255, cv2.THRESH_BINARY)
参数调整后输出:调参后二值化结果,预期二值化效果参考:预期效果
- 当前使用的完整答案检测代码:
def solve(img,n_row = 50): height, width, channels = img.shape n_col = 4 xShift = int(width/n_col) yShift = int(height/n_row) img = cv2.resize(img, (n_col * xShift, n_row*yShift)) img = cv2.GaussianBlur(img,(5,5),0) res, img = cv2.threshold(img, 60, 255, cv2.THRESH_BINARY) img = 255 - img for row in range(0, n_row): tmp_img = img [row*yShift + 5:(row+1)*yShift - 5,] area_sum = [] for col in range(n_col): area_sum.append(np.sum(tmp_img[1:,col*xShift :(col+1)*xShift])) y = str(area_sum > np.median(area_sum) * 1) result.append(area_sum > np.median(area_sum) * 5)
- 现有参考优化思路:统计每个边界矩形内的白色像素数量,过滤面积小于设定阈值的轮廓,参考代码:
inputImg= cv2.imread('input.jpg') img = cv2.cvtColor(inputImg, cv2.COLOR_BGR2GRAY) mask = np.zeros(img.shape[:2], dtype=img.dtype) ret, otsu_threshold = cv2.threshold(img, 120, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) contours, hierarchy = cv2.findContours(otsu_threshold, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE) for c in contours: x,y,w,h = cv2.boundingRect(c) if cv2.contourArea(c) > 1500: cv2.rectangle(otsu_threshold, (x, y), (x+w, y+h), (0,255,0), 2) cv2.imshow('Otsu', otsu_threshold) cv2.waitKey(0)
- 需求:给出可优化二值化效果、实现所有填涂答案准确检测的落地方案。
落地方案
1. 重构预处理流程解决二值化噪点问题
- 必须先将三通道BGR图像转为单通道灰度图,再做模糊、阈值操作,避免通道间数值差引入噪点。
- 弃用手动设置的固定阈值,改用OTSU全局自适应阈值或局部自适应阈值,自动适配图像光照、印刷色差,分割精度远高于固定阈值。
- 二值化后增加形态学开运算,去除孤立的小噪点,同时不会破坏填涂区域的连通性。
预处理参考代码:
import cv2 import numpy as np # 读入图像 img = cv2.imread('input.png') # 转单通道灰度图 gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 高斯模糊平滑噪点 blur = cv2.GaussianBlur(gray, (5,5), 0) # OTSU二值化,反转后填涂区域为白色、背景为黑色 _, thresh = cv2.threshold(blur, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU) # 3x3核做开运算,去除像素级小噪点 kernel = np.ones((3,3), np.uint8) clean_thresh = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=1)
2. 增加答题卡矫正与ROI提取步骤
当前直接按行列均分图像的逻辑对倾斜、偏移的答题卡容错率为0,需先通过轮廓检测找到答题卡最大外边框,做透视变换将答题卡矫正为正视图,再单独裁剪出答题区域作为检测ROI,排除边框、标题、考号区等无关内容的干扰。
3. 优化填涂检测逻辑
- 弃用网格均分统计像素的逻辑,先检测二值图的所有外轮廓,通过宽高比、面积区间筛选出选项气泡:标准OMR选项为近似正圆/正方形,宽高比在0.81.2区间,面积根据图像分辨率设置区间(如300DPI扫描图的气泡面积通常在100500像素范围),直接过滤线条、文字、噪点等无关轮廓。
- 将筛选出的气泡按坐标排序,逐行分组后统计每个气泡内的白色像素占比,占比超过设定阈值(通常为35%~45%)即判定为已填涂,稳定性远高于中位数乘系数的判断逻辑。
填涂检测参考代码片段:
# 检测外轮廓 contours, _ = cv2.findContours(clean_thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) bubbles = [] for c in contours: x,y,w,h = cv2.boundingRect(c) aspect_ratio = w / float(h) area = cv2.contourArea(c) # 过滤符合气泡尺寸、形状特征的轮廓 if 0.8 <= aspect_ratio <= 1.2 and 100 <= area <= 500: bubbles.append( (x,y,w,h) ) # 先按行分组、再按列排序所有气泡 bubbles = sorted(bubbles, key=lambda b: (b[1]//20, b[0])) # 逐题判断填涂选项 for q_idx in range(0, len(bubbles), 4): row_bubbles = bubbles[q_idx:q_idx+4] fill_ratios = [] for (x,y,w,h) in row_bubbles: bubble_roi = clean_thresh[y:y+h, x:x+w] # 计算气泡内填涂像素占比 fill_ratio = cv2.countNonZero(bubble_roi) / (w*h) fill_ratios.append(fill_ratio) # 取占比最高且超过阈值的选项作为填涂结果 selected = np.argmax(fill_ratios) if fill_ratios[selected] > 0.4: print(f"第{q_idx//4 +1}题答案:{chr(ord('A')+selected)}")
内容的提问来源于stack exchange,提问作者rnative
相关产品推荐
相关产品推荐

