You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用OpenCV去除矩形轮廓 为Pytesseract提取图像文本

问题描述

我想要提取一张图像中的文本,尝试去除矩形轮廓时,通过检测构成方框的水平线和垂直线来处理,但发现部分字符像素被误识别为垂直线,导致文本受损。希望得到不含矩形框、仅保留文本行的干净图像,以便后续用Pytesseract提取文本。

我的尝试代码

import cv2
from PIL import Image
import matplotlib.pylab as plt
import skimage.io as io  # 补充代码中缺失的IO库导入

image = io.imread("sample.png")
result = image.copy()
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]

# 去除水平线
horizontal_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (40,1))
remove_horizontal = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, horizontal_kernel, iterations=2)
cnts = cv2.findContours(remove_horizontal, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
cnts = cnts[0] if len(cnts) == 2 else cnts[1]
for c in cnts:
    cv2.drawContours(result, [c], -1, (255,255,255), 5)
plt.imshow(result)

# 去除垂直线
vertical_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1,40))
remove_vertical = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, vertical_kernel, iterations=2)
cnts = cv2.findContours(remove_vertical, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
cnts = cnts[0] if len(cnts) == 2 else cnts[1]
for c in cnts:
    cv2.drawContours(result, [c], -1, (255,255,255), 5)

plt.imshow(result)

改进方案

1. 调整形态学核尺寸

当前垂直线核高度(40)过大,易覆盖字符竖笔画。缩小核高度并减少迭代次数,只针对长框线处理:

# 调整垂直线核为(1,25),迭代1次
vertical_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1, 25))
remove_vertical = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, vertical_kernel, iterations=1)

2. 轮廓特征过滤框线

框线轮廓通常长宽比大、面积大,通过筛选轮廓特征区分框线与字符:

# 遍历垂直线轮廓,仅删除符合框线特征的部分
for c in cnts:
    x, y, w, h = cv2.boundingRect(c)
    # 根据图像实际尺寸调整阈值,保留短字符轮廓
    if h / w > 10 and h > 50:
        cv2.drawContours(result, [c], -1, (255,255,255), 5)

3. 霍夫直线检测精准去框

用霍夫变换检测长直线,仅删除符合框线特征的直线:

import numpy as np

# 边缘检测后检测直线
edges = cv2.Canny(gray, 50, 150)
lines = cv2.HoughLinesP(edges, 1, np.pi/180, threshold=100, minLineLength=100, maxLineGap=10)

# 用白色覆盖框线
for line in lines:
    x1, y1, x2, y2 = line[0]
    if abs(y2 - y1) < 5:  # 水平线
        cv2.line(result, (x1, y1), (x2, y2), (255,255,255), 3)
    elif abs(x2 - x1) < 5:  # 垂直线
        cv2.line(result, (x1, y1), (x2, y2), (255,255,255), 3)

4. 连通区域分析保留文本

通过连通区域的面积、宽高比筛选,仅保留字符区域:

import numpy as np

# 获取连通区域统计信息
num_labels, labels, stats, centroids = cv2.connectedComponentsWithStats(thresh, connectivity=8)

# 创建掩码保留字符区域
mask = np.zeros(gray.shape, dtype=np.uint8)
for i in range(1, num_labels):
    area = stats[i, cv2.CC_STAT_AREA]
    w = stats[i, cv2.CC_STAT_WIDTH]
    h = stats[i, cv2.CC_STAT_HEIGHT]
    # 根据字符尺寸调整阈值
    if 20 < area < 500 and 0.2 < w/h < 5:
        mask[labels == i] = 255

# 生成仅含文本的图像
text_only = cv2.bitwise_and(thresh, mask)
plt.imshow(text_only, cmap='gray')

内容的提问来源于stack exchange,提问作者lisa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 13:30:23