使用OpenCV+Python开发文档扫描仪时矩形边缘坐标偏移问题
文档扫描应用开发角点偏移问题排查
我正尝试基于OpenCV和Python从零开发一款文档扫描应用,目前已完成以下步骤:
- 图像重缩放
- 图像预处理:转换为灰度图、应用Gaussian blur、自适应阈值处理,最终执行Canny边缘检测
- 查找并绘制最大轮廓
- 检测并绘制轮廓边缘
当前问题出在第4步:检测得到的坐标点中有两个位置正确,但另外两个存在轻微偏移。我无法定位错误原因,同时想确认该问题是否可能由图像预处理的方式导致?
import cv2 import numpy as np # 图像重缩放函数 def Re_scaleImg(img): scale_percent = 50 width = int(img.shape[1] * scale_percent / 100) height = int(img.shape[0] * scale_percent / 100) dim = (width, height) # 重缩放图像 resized = cv2.resize(img, dim, interpolation = cv2.INTER_AREA) return resized # 图像处理函数 def process(img): gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) blur = cv2.GaussianBlur(gray, (3,3), 0) thresh = cv2.adaptiveThreshold(blur, 255, cv2.ADAPTIVE_THRESH_MEAN_C, cv2.THRESH_BINARY, 11, 2) edged = cv2.Canny(thresh, 75, 200) #cv2.imshow("blur", blur) #cv2.imshow("edged", thresh) return edged # 计算轮廓面积函数 def find_contourArea(contours): areas = [] for cnt in contours: cont_area = cv2.contourArea(cnt) areas.append(cont_area) return areas image = cv2.imread("receipt.jpeg") resized = Re_scaleImg(image) processed_img = process(resized) # 查找轮廓 contours, hierarchy = cv2.findContours(processed_img.copy(), cv2.RETR_LIST, cv2.CHAIN_APPROX_SIMPLE) resized_copy1 = resized.copy() # 轮廓按面积倒序排序 sorted_contours = sorted(contours, key=cv2.contourArea, reverse=True) largest_contour = sorted_contours[0] epsilon = 0.01*cv2.arcLength(largest_contour, True) approximation = cv2.approxPolyDP(largest_contour, epsilon, True) cv2.drawContours(resized_copy1, [approximation], -1, (0, 255, 0), 3) # 获取矩形角点 rot_rect = cv2.minAreaRect(largest_contour) box = cv2.boxPoints(rot_rect) box = np.int0(box) for p in box: pt = (p[0], p[1]) cv2.circle(resized_copy1, pt, 10, (255, 0, 0), -1) print(pt) cv2.imshow("contours", resized_copy1) cv2.waitKey(0)
相关图片
原始图像

输出图像

问题根因与修复方案
问题根因
角点偏移确实和预处理逻辑、轮廓处理逻辑直接相关,主要问题如下:
- 当前取角点用的
cv2.minAreaRect是基于检测到的轮廓拟合最小外接矩形,预处理后轮廓夹杂了收据表面的文字、阴影噪点,拟合出来的矩形自然会有偏移 - 预处理的高斯模糊核为(3,3),模糊程度不足以过滤收据表面的打印文字噪点,导致检测到的最大轮廓不是纯收据外轮廓,夹带了内层边缘
- 已经用
approxPolyDP拟合了四边形轮廓,不需要再调用minAreaRect多做一次拟合,额外步骤会引入不必要的误差 - 轮廓检索模式用了
RETR_LIST会检索所有层级的轮廓,内层文字轮廓会干扰最大轮廓的排序判断
修复方案
- 调整预处理参数:把高斯模糊核改成(5,5),适当提高模糊程度过滤表面噪点;自适应阈值改用高斯加权模式+反色,让边缘特征更突出;调整Canny高低阈值为50和150,避免边缘过碎
- 优先使用
approxPolyDP的拟合结果取角点:拟合后先判断结果长度是不是4,确认是四边形后直接用这四个点作为角点,跳过minAreaRect步骤 - 轮廓检索模式改成
RETR_EXTERNAL,只取最外层轮廓,排除内层文字轮廓的干扰
修改后的核心代码片段如下:
# 调整后的预处理函数 def process(img): gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) blur = cv2.GaussianBlur(gray, (5,5), 0) thresh = cv2.adaptiveThreshold(blur, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) edged = cv2.Canny(thresh, 50, 150) return edged # 轮廓检索改用RETR_EXTERNAL contours, hierarchy = cv2.findContours(processed_img.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # 取拟合后的四边形作为角点 epsilon = 0.02*cv2.arcLength(largest_contour, True) approximation = cv2.approxPolyDP(largest_contour, epsilon, True) if len(approximation) ==4: box = approximation.reshape(4,2) # 后续直接使用该box的四个点即可
内容的提问来源于stack exchange,提问作者Pranav Tyagi
相关产品推荐
相关产品推荐

