OpenCV裁剪图像所有字符时部分裁剪结果尺寸为零问题
问题描述
编写代码实现从图像中提取所有字符的功能,逻辑为识别字符轮廓后按从左到右排序,再将每个字符裁剪为独立图像。目前运行存在裁剪异常:并非所有字符都能被正确裁剪,部分字符裁剪后尺寸为零。

当前仅BCDEF五个字符裁剪后无零维度问题,输出效果如下:
完整实现代码如下:
import cv2 import numpy as np def crop_minAreaRect(img, rect): # 裁剪最小外接矩形对应区域 center = rect[0] size = rect[1] print("size[0]: " + str(int(size[0])) + ", size[1]: " + str(int(size[1]))) angle = rect[2] print("angle: " + str(angle)) rows,cols = img.shape[0], img.shape[1] M = cv2.getRotationMatrix2D((cols/2,rows/2),angle,1) img_rot = cv2.warpAffine(img,M,(cols,rows)) # 旋转 bounding box rect0 = (rect[0], rect[1], angle) box = cv2.boxPoints(rect0) pts = np.int0(cv2.transform(np.array([box]), M))[0] pts[pts < 0] = 0 # 裁剪 img_crop = img_rot[pts[1][1]:pts[0][1], pts[1][0]:pts[2][0]] w, h = img_crop.shape[0], img_crop.shape[1] print("w_cropped: " + str(w) + ", h_cropped: " + str(h)) return img_crop def sort_contours(cnts, method="left-to-right"): # 轮廓排序工具函数 reverse = False i = 0 if method == "right-to-left" or method == "bottom-to-top": reverse = True if method == "top-to-bottom" or method == "bottom-to-top": i = 1 boundingBoxes = [cv2.boundingRect(c) for c in cnts] (cnts, boundingBoxes) = zip(*sorted(zip(cnts, boundingBoxes), key=lambda b:b[1][i], reverse=reverse)) return (cnts, boundingBoxes) im_name = 'letters.png' im = cv2.imread(im_name) im_copy = im.copy() imgray = cv2.cvtColor(im, cv2.COLOR_BGR2GRAY) ret, thresh = cv2.threshold(imgray, 127, 255, 0) contours, hierarchy = cv2.findContours(thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) #cv2.drawContours(im_copy, contours, -1, (0,255,0), 2) #cv2.imshow("contours", im_copy) print("num contours: " + str(len(contours))) i = 0 sorted_cnts, bounding_boxes = sort_contours(contours, method="left-to-right") for cnt in sorted_cnts: size = cv2.contourArea(cnt) x,y,w,h = cv2.boundingRect(cnt) rect = cv2.minAreaRect(cnt) im_cropped = crop_minAreaRect(im, rect) h,w = im_cropped.shape[0], im_cropped.shape[1] if w > h: im_cropped = cv2.rotate(im_cropped, cv2.ROTATE_90_CLOCKWISE) print("w: " + str(w) + ", h: " + str(h)) if w>0 and h>0: cv2.imshow("cropped" + str(i), im_cropped) i += 1 cv2.waitKey(0)
故障原因
部分字符裁剪后尺寸为零是三个逻辑问题共同导致的:
- 旋转基准错误:旋转图像时使用整张图像的中心作为旋转原点,而非最小外接矩形自身的中心,旋转后矩形顶点的坐标映射完全错位,计算出的裁剪范围经常出现结束坐标小于起始坐标的情况,numpy切片遇到这种范围会直接返回空数组,对应尺寸为0的结果。
- 裁剪索引硬编码:固定用
pts[1][1]:pts[0][1]、pts[1][0]:pts[2][0]取裁剪范围,但cv2.boxPoints返回的四个顶点顺序不固定,旋转后顶点的上下、左右相对位置会发生变化,硬编码索引大概率取到反向的坐标范围。 - 二值化逻辑不匹配:测试图为白底黑字,代码中使用默认二值化模式会让黑色字符变为0值、白色背景变为255值,轮廓检测时部分字符会和背景融合,出现轮廓破碎、漏检的情况,进一步导致裁剪参数异常。
修复方法
针对以上问题做对应修改即可实现所有字符的正确裁剪:
- 调整二值化参数,使用反二值化模式将深色字符转为白色前景,保证每个字符的轮廓都能被完整检测
- 重写裁剪逻辑,旋转时以最小外接矩形自身中心为基准,不再硬编码顶点索引,直接通过矩形中心和宽高计算裁剪范围,同时增加坐标边界保护避免越界
- 增加轮廓面积过滤,剔除图像噪点生成的无效极小轮廓
修改后的核心代码片段如下:
# 1. 修改二值化逻辑 ret, thresh = cv2.threshold(imgray, 127, 255, cv2.THRESH_BINARY_INV) # 2. 重写裁剪函数 def crop_minAreaRect(img, rect): center, size, angle = rect target_w, target_h = int(size[0]), int(size[1]) # 以矩形自身中心为旋转基准 M = cv2.getRotationMatrix2D(center, angle, 1.0) img_rot = cv2.warpAffine(img, M, (img.shape[1], img.shape[0])) # 直接通过中心和宽高计算裁剪坐标 x_start = int(center[0] - target_w / 2) y_start = int(center[1] - target_h / 2) # 边界保护,避免坐标越界 x1 = max(0, x_start) y1 = max(0, y_start) x2 = min(img.shape[1], x_start + target_w) y2 = min(img.shape[0], y_start + target_h) return img_rot[y1:y2, x1:x2] # 3. 遍历轮廓时增加噪点过滤 for cnt in sorted_cnts: contour_area = cv2.contourArea(cnt) # 过滤面积小于10像素的无效噪点轮廓 if contour_area < 10: i += 1 continue rect = cv2.minAreaRect(cnt) im_cropped = crop_minAreaRect(im, rect) h_crop, w_crop = im_cropped.shape[:2] if w_crop > h_crop: im_cropped = cv2.rotate(im_cropped, cv2.ROTATE_90_CLOCKWISE) if w_crop > 0 and h_crop > 0: cv2.imshow("cropped" + str(i), im_cropped) i += 1
内容的提问来源于stack exchange,提问作者Nick_F
相关产品推荐
相关产品推荐

