如何用OpenCV去除图像中的线条且不损坏其余内容?
问题描述
我有一张带有字符和水平线条的图像,需要去除线条但完整保留所有字符。
首次尝试及问题
我先用以下代码尝试去除水平线条:
image = cv2.imread('im.png') gray = cv2.cvtColor(image,cv2.COLOR_BGR2GRAY) thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # Remove horizontal lines horizontal_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (20,1)) remove_horizontal = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, horizontal_kernel, iterations=2) cnts = cv2.findContours(remove_horizontal, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) cnts = cnts[0] if len(cnts) == 2 else cnts[1] for c in cnts: cv2.drawContours(result, [c], -1, (255,255,255), 5) plt.imshow(result)
处理后发现,和线条相连的字符部分也被一并删除了,这些字符是必须保留的。
二次尝试及残留问题
我又尝试添加修复逻辑,代码如下:
image = cv2.imread('connected/im.png') gray = cv2.cvtColor(image,cv2.COLOR_BGR2GRAY) thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # Remove horizontal horizontal_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (40,1)) detected_lines = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, horizontal_kernel, iterations=2) cnts = cv2.findContours(detected_lines, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) cnts = cnts[0] if len(cnts) == 2 else cnts[1] for c in cnts: cv2.drawContours(image, [c], -1, (255,255,255), 2) # Repair image repair_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3,6)) result = 255 - cv2.morphologyEx(255 - image, cv2.MORPH_CLOSE, repair_kernel, iterations=2) plt.imshow(result)
虽然修复了部分字符,但图像出现明显颗粒感,仍能看到线条残留的痕迹。
请问有没有更好的图像修复/线条去除方法?
优化方案
这里提供几种更精准的线条去除+字符修复方法,针对你的场景可以按需选择:
方法1:结合形态学检测与图像修复(Inpaint)
先精准定位线条区域,再用图像修复算法填充线条位置,参考周围像素还原字符:
import cv2 import numpy as np # 读取图像并预处理 image = cv2.imread('im.png') gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # 检测水平线条:用长核确保只捕获线条,不包含字符 horizontal_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (50, 1)) detected_lines = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, horizontal_kernel, iterations=1) # 生成修复掩码:标记需要修复的线条区域 mask = np.zeros(image.shape[:2], np.uint8) cnts = cv2.findContours(detected_lines, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) cnts = cnts[0] if len(cnts) == 2 else cnts[1] for c in cnts: # 控制线条绘制宽度,避免覆盖过多字符 cv2.drawContours(mask, [c], -1, 255, 2) # 使用Telea算法修复图像,适配纹理/结构类修复场景 result = cv2.inpaint(image, mask, 3, cv2.INPAINT_TELEA) cv2.imshow('Result', result) cv2.waitKey(0)
优势:精准区分线条与字符,修复时参考周边像素填充,不会破坏字符完整性,无颗粒感残留。
方法2:改进形态学流程,分离线条与字符
通过膨胀-分离-腐蚀的操作顺序,先让字符与线条断开连接,再去除线条:
import cv2 import numpy as np image = cv2.imread('im.png') gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # 先膨胀字符,让粘连的字符与线条分离 dilate_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2, 2)) dilated = cv2.dilate(thresh, dilate_kernel, iterations=1) # 检测线条,此时线条与字符已无粘连 horizontal_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (40, 1)) detected_lines = cv2.morphologyEx(dilated, cv2.MORPH_OPEN, horizontal_kernel, iterations=1) # 从膨胀后的图像中减去线条区域,再腐蚀回字符原始大小 result = cv2.subtract(dilated, detected_lines) erode_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2, 2)) result = cv2.erode(result, erode_kernel, iterations=1) # 转RGB格式显示 result = cv2.cvtColor(result, cv2.COLOR_GRAY2BGR) cv2.imshow('Result', result) cv2.waitKey(0)
优势:避免线条与字符粘连导致的误删,修复后的字符轮廓完整,无残留痕迹。
方法3:基于轮廓特征过滤,只删除线条
线条的轮廓通常是细长型,与字符的轮廓面积、宽高比差异明显,可通过轮廓特征区分:
import cv2 import numpy as np image = cv2.imread('im.png') gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1] # 提取所有轮廓 cnts = cv2.findContours(thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) cnts = cnts[0] if len(cnts) == 2 else cnts[1] # 创建掩码,保留字符轮廓,过滤线条轮廓 mask = np.zeros(thresh.shape, np.uint8) for c in cnts: area = cv2.contourArea(c) x, y, w, h = cv2.boundingRect(c) # 线条的宽高比远大于字符,且面积相对均匀 if w / h > 20 and area > 100: continue cv2.drawContours(mask, [c], -1, 255, -1) # 用掩码提取字符,背景填充白色 result = np.full(image.shape, 255, dtype=np.uint8) result[mask == 255] = image[mask == 255] cv2.imshow('Result', result) cv2.waitKey(0)
优势:完全基于轮廓特征区分线条和字符,适合线条与字符形态差异明显的场景,处理效果干净彻底。
内容的提问来源于stack exchange,提问作者InterestingQuestions44
相关产品推荐
相关产品推荐

