如何在Python中通过边缘检测分割numpy数组中的字符图像
问题
我的图像存储在numpy数组中,我希望将这些图像分割为包含单个字符的独立图像。
输入图像:
我尝试了以下代码:
import cv2 # 加载图像、灰度化、高斯模糊、大津阈值处理、膨胀操作 image = arr # numpy_arr_containing 200X200 image original = image.copy() gray = image blur = cv2.GaussianBlur(gray, (5,5), 0) thresh = cv2.threshold(blur, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (15,15)) dilate = cv2.dilate(thresh, kernel, iterations=2) # 查找轮廓、获取边界框坐标并提取ROI cnts = cv2.findContours(dilate, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) cnts = cnts[0] if len(cnts) == 2 else cnts[1] image_number = 0 for c in cnts: x,y,w,h = cv2.boundingRect(c) cv2.rectangle(image, (x, y), (x + w, y + h), (36,255,12), 3) ROI = original[y:y+h, x:x+w] cv2.imwrite("ROI_{}.png".format(image_number), ROI) image_number += 1 # cv2.imshow('image', image) # cv2.imshow('thresh', thresh) # cv2.imshow('dilate', dilate) # cv2.waitKey()
解决方案
你的代码核心思路可行,但当前使用的膨胀核尺寸(15,15)过大,容易导致相邻字符的轮廓被合并,无法分割出单个字符。以下是优化后的实现:
优化代码
import cv2 import numpy as np # 假设arr是你的200x200 numpy数组图像 image = arr original = image.copy() # 兼容彩色/灰度图像的灰度化处理 if len(image.shape) == 3: gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) else: gray = image # 高斯模糊+大津阈值二值化 blur = cv2.GaussianBlur(gray, (3, 3), 0) thresh = cv2.threshold(blur, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # 使用小尺寸膨胀核,避免字符轮廓合并 kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3, 3)) dilate = cv2.dilate(thresh, kernel, iterations=1) # 提取外部轮廓 cnts = cv2.findContours(dilate, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) cnts = cnts[0] if len(cnts) == 2 else cnts[1] # 按x坐标排序轮廓,保证字符顺序与原图一致 cnts = sorted(cnts, key=lambda x: cv2.boundingRect(x)[0]) image_number = 0 for c in cnts: x, y, w, h = cv2.boundingRect(c) # 过滤过小的噪点轮廓 if w > 10 and h > 10: cv2.rectangle(image, (x, y), (x + w, y + h), (36, 255, 12), 2) ROI = original[y:y+h, x:x+w] cv2.imwrite(f"ROI_{image_number}.png", ROI) image_number += 1 # 可选:显示结果窗口 # cv2.imshow('分割结果', image) # cv2.imshow('二值化图', thresh) # cv2.waitKey(0) # cv2.destroyAllWindows()
关键优化点
- 调整膨胀核尺寸:将大核改为
(3,3),避免相邻字符轮廓被合并。 - 轮廓排序:按x坐标排序轮廓,保证分割出的字符顺序与原图一致。
- 噪点过滤:通过判断轮廓宽高,过滤掉无关噪点,只保留有效字符区域。
- 灰度化兼容:增加对彩色图像的处理逻辑,提升代码通用性。
内容的提问来源于stack exchange,提问作者Niteesh Chowdary
相关产品推荐
相关产品推荐

