You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

印地语图像ROI检测问题:两行粗体文本被同一框识别需拆分

印地语文本行ROI检测问题

我尝试用ROI框对印地语图像中的每一行文本做精准检测,但现在遇到个问题:两行大粗体文本被识别进了同一个ROI里,效果如下:
检测结果
原始图像:
原始图像

需要把每一行文本都准确检测为独立的ROI,以下是我目前使用的源代码:

import cv2
from google.colab.patches import cv2_imshow
import numpy as np

if __name__ == "__main__":
  image = cv2.imread('datasets/0010_jpg.rf.e7741188a2afa6db3dee4324e8486a34.jpg')

  # Display the image
  # cv2_imshow(image)

  # Convert image to grayscale
  gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
  # cv2_imshow(gray)

  # Convert grayscale image to binary
  ret, thresh = cv2.threshold(gray, 150, 255, cv2.THRESH_BINARY_INV)
  # cv2_imshow(thresh)

   # Apply Canny edge detection
  edges = cv2.Canny(thresh, 50, 150)  # Adjust the threshold values as needed
  # cv2_imshow(edges)

  # Dilation
  kernel = np.ones((5, 200), np.uint8)
  img_dilation = cv2.dilate(edges, kernel, iterations=1)
  # cv2_imshow(img_dilation)

  # Find contours
  contours, hierarchy = cv2.findContours(img_dilation.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)

  # Sort contours based on their bounding box coordinates
  bounding_boxes = [cv2.boundingRect(ctr) for ctr in contours]
  sorted_contours = [ctr for _, ctr in sorted(zip(bounding_boxes, contours), key=lambda pair: pair[0][1])]

  # Loop over sorted contours
  for i, ctr in enumerate(sorted_contours):
      # Get bounding box
      x, y, w, h = cv2.boundingRect(ctr)

      # Getting ROI
      roi = image[y:y+h-5, x:x+w]
      roi_row = roi.shape[0]
      roi_col = roi.shape[1]

      # Show ROI
      if(roi_row>3000 or roi_row<=20 or roi_row<=10 or roi_col<=110):
          continue
      print(i)
      print(roi_row,roi_col)
      cv2_imshow(roi)
      cv2.rectangle(image, (x, y), (x + w, y + h), (90, 0, 255), 2)

  cv2_imshow(image)

问题原因及解决方法

当前代码里的**膨胀核尺寸(5,200)**是横向过大的矩形,会把上下相邻的粗体文本区域连在一起,导致合并成一个ROI。另外,Canny边缘检测不是必须的步骤,反而可能干扰文本区域的连通性判断。

可以按以下步骤修改代码:

1. 调整膨胀核,避免上下行合并

把膨胀核改成纵向小、横向适中的尺寸,比如(2, 100),这样既能把同一行的文本连起来,又不会把上下两行的粗体文本合并。

2. 跳过边缘检测,直接处理二值图

取消Canny边缘检测步骤,直接对二值化后的图像做膨胀处理,更准确保留文本区域的连通性。

3. 简化ROI过滤条件

原代码里的roi_row<=20 or roi_row<=10属于重复判断,简化为清晰的阈值过滤逻辑。

修改后的完整代码

import cv2
from google.colab.patches import cv2_imshow
import numpy as np

if __name__ == "__main__":
    image = cv2.imread('datasets/0010_jpg.rf.e7741188a2afa6db3dee4324e8486a34.jpg')

    # 转灰度图
    gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

    # 二值化(反相,让文本为白色)
    ret, thresh = cv2.threshold(gray, 150, 255, cv2.THRESH_BINARY_INV)

    # 调整膨胀核:纵向小尺寸,避免合并上下行
    kernel = np.ones((2, 100), np.uint8)
    img_dilation = cv2.dilate(thresh, kernel, iterations=1)

    # 查找轮廓
    contours, hierarchy = cv2.findContours(img_dilation.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)

    # 按y坐标排序轮廓(从上到下)
    bounding_boxes = [cv2.boundingRect(ctr) for ctr in contours]
    sorted_contours = [ctr for _, ctr in sorted(zip(bounding_boxes, contours), key=lambda pair: pair[0][1])]

    # 遍历处理每个轮廓
    for i, ctr in enumerate(sorted_contours):
        x, y, w, h = cv2.boundingRect(ctr)
        roi = image[y:y+h-5, x:x+w]
        roi_row, roi_col = roi.shape[:2]

        # 过滤无效ROI:排除过大或过小的区域
        if roi_row > 3000 or roi_row <= 20 or roi_col <= 110:
            continue
        
        print(f"ROI {i}: 尺寸 {roi_row}x{roi_col}")
        cv2_imshow(roi)
        cv2.rectangle(image, (x, y), (x + w, y + h), (90, 0, 255), 2)

    cv2_imshow(image)

额外优化建议

如果仍存在个别行合并的情况,可以尝试:

  • 调整二值化阈值(比如把150改成120-180之间的值),让文本和背景分离更彻底
  • 对膨胀后的图像做一次小尺寸腐蚀,消除行与行之间的细微连接
  • 使用cv2.RETR_LIST代替cv2.RETR_EXTERNAL,配合轮廓层级信息过滤嵌套区域

内容的提问来源于stack exchange,提问作者Anish Khatiwada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 08:10:31