You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用OpenCV+Python开发文档扫描仪时矩形边缘坐标偏移问题

文档扫描应用开发角点偏移问题排查

我正尝试基于OpenCV和Python从零开发一款文档扫描应用,目前已完成以下步骤:

  • 图像重缩放
  • 图像预处理:转换为灰度图、应用Gaussian blur、自适应阈值处理,最终执行Canny边缘检测
  • 查找并绘制最大轮廓
  • 检测并绘制轮廓边缘

当前问题出在第4步:检测得到的坐标点中有两个位置正确,但另外两个存在轻微偏移。我无法定位错误原因,同时想确认该问题是否可能由图像预处理的方式导致?

import cv2
import numpy as np

# 图像重缩放函数
def Re_scaleImg(img):
    scale_percent = 50
    width = int(img.shape[1] * scale_percent / 100)
    height = int(img.shape[0] * scale_percent / 100)
    dim = (width, height)

    # 重缩放图像
    resized = cv2.resize(img, dim, interpolation = cv2.INTER_AREA)
    return resized


# 图像处理函数
def process(img):
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    blur = cv2.GaussianBlur(gray, (3,3), 0)
    thresh = cv2.adaptiveThreshold(blur, 255, cv2.ADAPTIVE_THRESH_MEAN_C, cv2.THRESH_BINARY, 11, 2)
    edged = cv2.Canny(thresh, 75, 200)
    #cv2.imshow("blur", blur)
    #cv2.imshow("edged", thresh)
    return edged

# 计算轮廓面积函数
def find_contourArea(contours):
    areas = []
    for cnt in contours:
        cont_area = cv2.contourArea(cnt)
        areas.append(cont_area)
    return areas


image = cv2.imread("receipt.jpeg")
resized = Re_scaleImg(image)
processed_img = process(resized)

# 查找轮廓
contours, hierarchy = cv2.findContours(processed_img.copy(), cv2.RETR_LIST, cv2.CHAIN_APPROX_SIMPLE)
resized_copy1 = resized.copy()

# 轮廓按面积倒序排序
sorted_contours = sorted(contours, key=cv2.contourArea, reverse=True)

largest_contour = sorted_contours[0]
epsilon = 0.01*cv2.arcLength(largest_contour, True)
approximation = cv2.approxPolyDP(largest_contour, epsilon, True)

cv2.drawContours(resized_copy1, [approximation], -1, (0, 255, 0), 3)

# 获取矩形角点
rot_rect = cv2.minAreaRect(largest_contour)
box = cv2.boxPoints(rot_rect)
box = np.int0(box)
for p in box:
    pt = (p[0], p[1])
    cv2.circle(resized_copy1, pt, 10, (255, 0, 0), -1)
    print(pt)

cv2.imshow("contours", resized_copy1)

cv2.waitKey(0)

相关图片

原始图像

原始图像

输出图像

输出图像

问题根因与修复方案

问题根因

角点偏移确实和预处理逻辑、轮廓处理逻辑直接相关,主要问题如下:

  1. 当前取角点用的cv2.minAreaRect是基于检测到的轮廓拟合最小外接矩形,预处理后轮廓夹杂了收据表面的文字、阴影噪点,拟合出来的矩形自然会有偏移
  2. 预处理的高斯模糊核为(3,3),模糊程度不足以过滤收据表面的打印文字噪点,导致检测到的最大轮廓不是纯收据外轮廓,夹带了内层边缘
  3. 已经用approxPolyDP拟合了四边形轮廓,不需要再调用minAreaRect多做一次拟合,额外步骤会引入不必要的误差
  4. 轮廓检索模式用了RETR_LIST会检索所有层级的轮廓,内层文字轮廓会干扰最大轮廓的排序判断

修复方案

  1. 调整预处理参数:把高斯模糊核改成(5,5),适当提高模糊程度过滤表面噪点;自适应阈值改用高斯加权模式+反色,让边缘特征更突出;调整Canny高低阈值为50和150,避免边缘过碎
  2. 优先使用approxPolyDP的拟合结果取角点:拟合后先判断结果长度是不是4,确认是四边形后直接用这四个点作为角点,跳过minAreaRect步骤
  3. 轮廓检索模式改成RETR_EXTERNAL,只取最外层轮廓,排除内层文字轮廓的干扰

修改后的核心代码片段如下:

# 调整后的预处理函数
def process(img):
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    blur = cv2.GaussianBlur(gray, (5,5), 0) 
    thresh = cv2.adaptiveThreshold(blur, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    edged = cv2.Canny(thresh, 50, 150)
    return edged

# 轮廓检索改用RETR_EXTERNAL
contours, hierarchy = cv2.findContours(processed_img.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)

# 取拟合后的四边形作为角点
epsilon = 0.02*cv2.arcLength(largest_contour, True)
approximation = cv2.approxPolyDP(largest_contour, epsilon, True)
if len(approximation) ==4:
    box = approximation.reshape(4,2)
    # 后续直接使用该box的四个点即可

内容的提问来源于stack exchange,提问作者Pranav Tyagi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 22:09:04