You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python和OpenCV3识别流程图中的文本区域?

优化OpenCV3文本区域轮廓检测的方案

嘿,我瞅了你这段用OpenCV 3做文本区域检测的代码,目前只能识别部分轮廓是吧?咱们来捋捋问题出在哪,再给你优化下方案。

你提供的现有代码(补全未完成部分)

import os,sys,cv2,pytesseract

## IMAGE
afile = "test-small.jpg"
def reader(afile):
    aimg = cv2.imread(afile,0)
    print("Image Shape%s | Size:%s" % (aimg.shape,aimg.size))
    return aimg

def boundbox(aimg):
    out_path2 = "%s-tagged.jpg" % (afile.split('.')[0])
    # 原代码未完成的轮廓检测逻辑

问题分析与优化步骤

你的代码目前只做了灰度图读取,缺少关键的预处理步骤和合理的轮廓检测参数,这就是为啥只能检测到部分轮廓的原因。咱们一步步优化:

  • 第一步:增强预处理,突出文本轮廓
    灰度图之后,必须做二值化和形态学操作,把文本和背景彻底分离,还能把断开的文本区域连起来。推荐用Otsu自动阈值二值化,不用手动调参数,适配性更强。

    # 改造reader函数,同时返回彩色图(用来画框)和灰度图
    def reader(afile):
        aimg_color = cv2.imread(afile)
        aimg_gray = cv2.cvtColor(aimg_color, cv2.COLOR_BGR2GRAY)
        print(f"Image Shape {aimg_gray.shape} | Size: {aimg_gray.size}")
        return aimg_color, aimg_gray
    
    # 预处理核心逻辑
    _, thresh_img = cv2.threshold(aimg_gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
    # 用小矩形核做膨胀,连接断开的文本笔画
    kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2))
    dilated_img = cv2.dilate(thresh_img, kernel, iterations=1)
    
  • 第二步:调整轮廓检测参数,过滤噪声
    用cv2.RETR_EXTERNAL只检测最外层轮廓,避免把文本内部的空隙也当成轮廓;用cv2.CHAIN_APPROX_SIMPLE压缩轮廓点,减少计算量。还要过滤掉太小的轮廓,避免把噪声当成文本。

    def boundbox(aimg_color, aimg_gray):
        # 预处理
        _, thresh = cv2.threshold(aimg_gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
        kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2))
        dilated = cv2.dilate(thresh, kernel, iterations=1)
        # 检测最外层轮廓
        contours, _ = cv2.findContours(dilated, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
        # 遍历轮廓画矩形并过滤小噪声
        for cnt in contours:
            x, y, w, h = cv2.boundingRect(cnt)
            # 根据你的图像实际大小调整阈值,过滤无效小区域
            if w > 15 and h > 15:
                cv2.rectangle(aimg_color, (x,y), (x+w, y+h), (0,255,0), 2)
        # 保存结果
        out_path = f"{afile.split('.')[0]}-tagged.jpg"
        cv2.imwrite(out_path, aimg_color)
        print(f"标注后的图像已保存到 {out_path}")
    

完整可运行代码

把上面的部分整合起来,就是完整的解决方案:

import os, sys, cv2, pytesseract

def reader(afile):
    aimg_color = cv2.imread(afile)
    aimg_gray = cv2.cvtColor(aimg_color, cv2.COLOR_BGR2GRAY)
    print(f"Image Shape {aimg_gray.shape} | Size: {aimg_gray.size}")
    return aimg_color, aimg_gray

def boundbox(aimg_color, aimg_gray):
    # 二值化(反色处理,让文本变成白色,背景黑色)
    _, thresh = cv2.threshold(aimg_gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
    # 形态学膨胀,连接断开的文本区域
    kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2,2))
    dilated = cv2.dilate(thresh, kernel, iterations=1)
    # 检测最外层轮廓
    contours, _ = cv2.findContours(dilated, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    # 遍历画框并过滤小轮廓
    for cnt in contours:
        x, y, w, h = cv2.boundingRect(cnt)
        # 可根据你的图像实际大小调整这个阈值
        if w > 15 and h > 15:
            cv2.rectangle(aimg_color, (x, y), (x+w, y+h), (0, 255, 0), 2)
    # 保存标注后的图像
    out_path = f"{afile.split('.')[0]}-tagged.jpg"
    cv2.imwrite(out_path, aimg_color)
    print(f"标注完成,结果已保存至 {out_path}")

if __name__ == "__main__":
    afile = "test-small.jpg"
    color_img, gray_img = reader(afile)
    boundbox(color_img, gray_img)

关键说明

  • Otsu阈值二值化:自动计算最优阈值,不用手动调整,适配不同光照条件的图像
  • 形态学膨胀:解决文本笔画断开导致的轮廓碎片化问题
  • 轮廓过滤:排除小噪声点,避免标注无效区域
  • 彩色图绘制:灰度图无法显示彩色矩形框,所以用彩色原图来画框,确保标注清晰可见

内容的提问来源于stack exchange,提问作者Bade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:18:40