印地语图像ROI检测问题:两行粗体文本被同一框识别需拆分
印地语文本行ROI检测问题
我尝试用ROI框对印地语图像中的每一行文本做精准检测,但现在遇到个问题:两行大粗体文本被识别进了同一个ROI里,效果如下:
原始图像:
需要把每一行文本都准确检测为独立的ROI,以下是我目前使用的源代码:
import cv2 from google.colab.patches import cv2_imshow import numpy as np if __name__ == "__main__": image = cv2.imread('datasets/0010_jpg.rf.e7741188a2afa6db3dee4324e8486a34.jpg') # Display the image # cv2_imshow(image) # Convert image to grayscale gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) # cv2_imshow(gray) # Convert grayscale image to binary ret, thresh = cv2.threshold(gray, 150, 255, cv2.THRESH_BINARY_INV) # cv2_imshow(thresh) # Apply Canny edge detection edges = cv2.Canny(thresh, 50, 150) # Adjust the threshold values as needed # cv2_imshow(edges) # Dilation kernel = np.ones((5, 200), np.uint8) img_dilation = cv2.dilate(edges, kernel, iterations=1) # cv2_imshow(img_dilation) # Find contours contours, hierarchy = cv2.findContours(img_dilation.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # Sort contours based on their bounding box coordinates bounding_boxes = [cv2.boundingRect(ctr) for ctr in contours] sorted_contours = [ctr for _, ctr in sorted(zip(bounding_boxes, contours), key=lambda pair: pair[0][1])] # Loop over sorted contours for i, ctr in enumerate(sorted_contours): # Get bounding box x, y, w, h = cv2.boundingRect(ctr) # Getting ROI roi = image[y:y+h-5, x:x+w] roi_row = roi.shape[0] roi_col = roi.shape[1] # Show ROI if(roi_row>3000 or roi_row<=20 or roi_row<=10 or roi_col<=110): continue print(i) print(roi_row,roi_col) cv2_imshow(roi) cv2.rectangle(image, (x, y), (x + w, y + h), (90, 0, 255), 2) cv2_imshow(image)
问题原因及解决方法
当前代码里的**膨胀核尺寸(5,200)**是横向过大的矩形,会把上下相邻的粗体文本区域连在一起,导致合并成一个ROI。另外,Canny边缘检测不是必须的步骤,反而可能干扰文本区域的连通性判断。
可以按以下步骤修改代码:
1. 调整膨胀核,避免上下行合并
把膨胀核改成纵向小、横向适中的尺寸,比如(2, 100),这样既能把同一行的文本连起来,又不会把上下两行的粗体文本合并。
2. 跳过边缘检测,直接处理二值图
取消Canny边缘检测步骤,直接对二值化后的图像做膨胀处理,更准确保留文本区域的连通性。
3. 简化ROI过滤条件
原代码里的roi_row<=20 or roi_row<=10属于重复判断,简化为清晰的阈值过滤逻辑。
修改后的完整代码
import cv2 from google.colab.patches import cv2_imshow import numpy as np if __name__ == "__main__": image = cv2.imread('datasets/0010_jpg.rf.e7741188a2afa6db3dee4324e8486a34.jpg') # 转灰度图 gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) # 二值化(反相,让文本为白色) ret, thresh = cv2.threshold(gray, 150, 255, cv2.THRESH_BINARY_INV) # 调整膨胀核:纵向小尺寸,避免合并上下行 kernel = np.ones((2, 100), np.uint8) img_dilation = cv2.dilate(thresh, kernel, iterations=1) # 查找轮廓 contours, hierarchy = cv2.findContours(img_dilation.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # 按y坐标排序轮廓(从上到下) bounding_boxes = [cv2.boundingRect(ctr) for ctr in contours] sorted_contours = [ctr for _, ctr in sorted(zip(bounding_boxes, contours), key=lambda pair: pair[0][1])] # 遍历处理每个轮廓 for i, ctr in enumerate(sorted_contours): x, y, w, h = cv2.boundingRect(ctr) roi = image[y:y+h-5, x:x+w] roi_row, roi_col = roi.shape[:2] # 过滤无效ROI:排除过大或过小的区域 if roi_row > 3000 or roi_row <= 20 or roi_col <= 110: continue print(f"ROI {i}: 尺寸 {roi_row}x{roi_col}") cv2_imshow(roi) cv2.rectangle(image, (x, y), (x + w, y + h), (90, 0, 255), 2) cv2_imshow(image)
额外优化建议
如果仍存在个别行合并的情况,可以尝试:
- 调整二值化阈值(比如把150改成120-180之间的值),让文本和背景分离更彻底
- 对膨胀后的图像做一次小尺寸腐蚀,消除行与行之间的细微连接
- 使用
cv2.RETR_LIST代替cv2.RETR_EXTERNAL,配合轮廓层级信息过滤嵌套区域
内容的提问来源于stack exchange,提问作者Anish Khatiwada
相关产品推荐
相关产品推荐

