如何将带曲线字母的凸包近似为四边形用于透视校正?
问题:带曲线字母的凸包四边形近似失败,如何实现预期校正?
背景与问题
输入图像:
预处理后图像:
该图像存在轻微倾斜,希望用cv2.warpPerspective()校正,尝试通过cv2.approxPolyDP()将凸包近似为四边形,核心代码片段:
hull = cv2.convexHull(contour) # 近似为四边形 epsilon = 0.1 * cv2.arcLength(hull, True) approximated = cv2.approxPolyDP(hull, epsilon, True)
预处理图像的凸包:
但近似结果不符合预期:
期望得到的四边形:
成功对比示例
无曲线字母的原图像:
其近似后得到正确四边形:
简言之:带曲线字母的凸包无法被正确近似为四边形,如何解决这个问题?
完整代码
def _perspective_fix_convex_hull(self): # RETR_TREE 提取所有轮廓并重建嵌套层次,用于保证精度 contours, _ = cv2.findContours(self.preprocessed_image, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE) # 绘制所有轮廓到新图像 self.image_contours = cv2.cvtColor(self.preprocessed_image, cv2.COLOR_GRAY2BGR) cv2.drawContours(self.image_contours, contours, -1, (0, 255, 0), 2) self.convex_hulls_contours = [] self.approximated_contours = [] for contour in contours: # 计算轮廓面积 area = cv2.contourArea(contour) if area < 200: # 过小的轮廓忽略,非目标字母 continue image = cv2.cvtColor(self.preprocessed_image, cv2.COLOR_GRAY2BGR) hull = cv2.convexHull(contour) cv2.drawContours(image, [hull], -1, (0, 255, 0), 2) self.convex_hulls_contours.append(image) # 近似为四边形 epsilon = 0.1 * cv2.arcLength(hull, True) approximated = cv2.approxPolyDP(hull, epsilon, True) # 绘制近似后的轮廓 image = cv2.cvtColor(self.preprocessed_image, cv2.COLOR_GRAY2BGR) cv2.drawContours(image, [approximated], -1, (255, 0, 0), 1) self.approximated_contours.append(image) if len(approximated) == 4: image = cv2.cvtColor(self.preprocessed_image, cv2.COLOR_GRAY2BGR) cv2.drawContours(image, [approximated], -1, (0, 0, 255), 2) self.final_contour = image letter_hull = [a[0] for a in approximated] # 去除多余维度 [[x, y]] -> [x, y] break # 创建矩形点占位符 rectangle = np.zeros((4, 2), dtype="float32") # 左上角点的x+y值最小,右下角点的x+y值最大 s = np.sum(letter_hull, axis=1) rectangle[0] = letter_hull[np.argmin(s)] rectangle[2] = letter_hull[np.argmax(s)] # 右上角点的x-y值最小,左下角点的x-y值最大 d = np.diff(letter_hull, axis=1) rectangle[1] = letter_hull[np.argmin(d)] rectangle[3] = letter_hull[np.argmax(d)] # 根据点计算新图像的宽高 (top_left, top_right, bottom_right, bottom_left) = rectangle # 取上下两边宽度的最大值作为最终宽度(勾股定理计算) width_top = np.sqrt(((top_right[0] - top_left[0]) ** 2) + ((top_right[1] - top_left[1]) ** 2)) width_bottom = np.sqrt(((bottom_right[0] - bottom_left[0]) ** 2) + ((bottom_right[1] - bottom_left[1]) ** 2)) width = max(int(width_top), int(width_bottom)) # 取左右两边高度的最大值作为最终高度(勾股定理计算) height_right = np.sqrt(((top_right[0] - bottom_right[0]) ** 2) + ((top_right[1] - bottom_right[1]) ** 2)) height_left = np.sqrt(((top_left[0] - bottom_left[0]) ** 2) + ((top_left[1] - bottom_left[1]) ** 2)) height = max(int(height_right), int(height_left)) # 创建目标点 destination = np.array([ [0, 0], [width - 1, 0], [width - 1, height - 1], [0, height - 1] ], dtype="float32") # 计算透视变换矩阵 matrix = cv2.getPerspectiveTransform(rectangle, destination) warped = cv2.warpPerspective(self.preprocessed_image, matrix, (width, height)) # 添加 padding self.warped_image = cv2.copyMakeBorder(warped, 10, 10, 10, 10, cv2.BORDER_CONSTANT, value=(0, 0, 0))
解决方案
方法1:动态调整epsilon阈值,强制筛选四边形
曲线字母的凸包轮廓点密度更高,固定比例的epsilon会保留过多曲线细节。可以动态调整epsilon值,直到得到四边形:
# 替换原近似代码段 epsilon_ratio = 0.05 # 初始更小的比例 max_attempts = 10 approximated = None for _ in range(max_attempts): epsilon = epsilon_ratio * cv2.arcLength(hull, True) approximated = cv2.approxPolyDP(hull, epsilon, True) if len(approximated) == 4: break epsilon_ratio += 0.02 # 逐步增大比例
方法2:用最小外接矩形替代凸包近似
绕开凸包的曲线细节问题,直接对凸包计算最小外接旋转矩形,提取四个顶点:
# 替换凸包近似部分 rect = cv2.minAreaRect(hull) box = cv2.boxPoints(rect) approximated = np.int0(box)
这种方法不受曲线影响,能直接得到规整的矩形顶点,适配带曲线的字符场景。
方法3:先膨胀预处理图像,简化轮廓
对预处理图像做膨胀操作,填充曲线的凹陷区域,让凸包更接近矩形,再进行近似:
# 在findContours前添加膨胀步骤 kernel = np.ones((3,3), np.uint8) dilated_img = cv2.dilate(self.preprocessed_image, kernel, iterations=1) contours, _ = cv2.findContours(dilated_img, cv2.RETR_TREE, cv2.CHAIN_APPROX_SIMPLE)
膨胀会让曲线部分的轮廓更趋近于矩形,后续的凸包近似更容易得到四边形。
内容的提问来源于stack exchange,提问作者R1D3R175
相关产品推荐
相关产品推荐

