You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免INPAINTING产生色块并优化图像文字去除效果?

图像Inpainting去除文字:色块问题修复与预处理优化方案

一、色块问题的核心原因与解决方法

出现色块的本质是掩码区域不准确或Inpainting算法参数适配性差:

  • 掩码误覆盖文字周边正常区域,导致算法填充时引入错误颜色信息
  • 固定Inpainting半径(当前为7)无法适配不同图片的文字大小与背景复杂度

具体修复措施:

  1. 精准过滤有效轮廓
    对检测到的轮廓做尺寸筛选,只保留符合文字宽高比、面积特征的轮廓,避免将背景噪点误判为文字区域:
    valid_cnts = []
    for c in cnts:
        x, y, w, h = cv2.boundingRect(c)
        # 文字宽高比通常在2-10区间,面积不小于50(可根据实际场景调整)
        if (2 < w/h < 10) and (w*h > 50):
            valid_cnts.append(c)
    
  2. 切换Inpainting算法并动态调整半径
    • 改用cv2.INPAINT_NS算法,它在纹理连续性上的表现优于TELEA,能有效减少色块
    • 根据文字平均宽度动态设置Inpainting半径,比如取轮廓平均宽度的1.2倍,避免固定半径适配性差的问题
  3. 放弃多次重复Inpainting操作
    多次重复操作会累积填充误差,加重色块问题,优先优化掩码精度后执行单次高质量Inpainting

二、针对白色文字的预处理优化

当前用THRESH_BINARY_INV+OTSU阈值对白色文字效果差,因为白色文字在灰度图中亮度高,反转阈值无法有效分离文字与背景,优化方案如下:

1. 动态切换阈值类型

根据图像灰度均值判断背景亮度,自动选择阈值模式:

gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
gray_mean = gray.mean()
if gray_mean < 127:
    # 背景偏暗,文字大概率为白色,用普通二值化
    ret, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU | cv2.THRESH_BINARY)
else:
    # 背景偏亮,文字偏暗,用反二值化
    ret, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU | cv2.THRESH_BINARY_INV)

2. 优化形态学操作流程

先腐蚀去除噪点,再膨胀连接文字笔画,同时调整核尺寸避免过度覆盖背景:

rect_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (12, 2))
# 先腐蚀去噪
erosion = cv2.erode(thresh, rect_kernel, iterations=1)
# 再膨胀连接文字笔画
dilation = cv2.dilate(erosion, rect_kernel, iterations=1)

3. 结合OCR生成精准掩码

若单纯靠阈值与轮廓效果不稳定,引入OCR工具(如Tesseract、PaddleOCR)直接定位文字区域,生成更精准的掩码:

from paddleocr import PaddleOCR

ocr = PaddleOCR(use_angle_cls=True, lang='en')
result = ocr.ocr(img, cls=True)

mask = np.ones(img.shape[:2], dtype="uint8") * 255
for line in result:
    for word_info in line:
        # 获取文字外接矩形坐标
        x1, y1 = int(word_info[0][0][0]), int(word_info[0][0][1])
        x2, y2 = int(word_info[0][2][0]), int(word_info[0][2][1])
        # 在掩码上标记文字区域(黑色为Inpainting目标区域)
        cv2.rectangle(mask, (x1, y1), (x2, y2), 0, -1)

三、完整优化后代码示例

import cv2
import numpy as np
import imutils

def preprocess(img):
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    # 动态选择阈值类型
    gray_mean = gray.mean()
    if gray_mean < 127:
        ret, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU | cv2.THRESH_BINARY)
    else:
        ret, thresh = cv2.threshold(gray, 0, 255, cv2.THRESH_OTSU | cv2.THRESH_BINARY_INV)
    
    # 优化形态学操作
    rect_kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (12, 2))
    erosion = cv2.erode(thresh, rect_kernel, iterations=1)
    dilation = cv2.dilate(erosion, rect_kernel, iterations=1)
    
    edged = cv2.Canny(dilation, 50, 100)
    cnts = cv2.findContours(edged.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    cnts = imutils.grab_contours(cnts)
    
    # 过滤有效轮廓
    valid_cnts = []
    for c in cnts:
        x, y, w, h = cv2.boundingRect(c)
        if (2 < w/h < 10) and (w*h > 50):
            valid_cnts.append(c)
    
    mask = np.ones(img.shape[:2], dtype="uint8") * 255
    for c in valid_cnts:
        cv2.drawContours(mask, [c], -1, 0, -1)
    return mask

# 加载图像
img = cv2.imread("input.jpg")
# 生成精准掩码
mask = preprocess(img)
# 执行Inpainting
inpaint_radius = 5  # 可根据文字大小调整
result = cv2.inpaint(img, mask, inpaint_radius, cv2.INPAINT_NS)
# 保存结果
cv2.imwrite("output.jpg", result)

内容的提问来源于stack exchange,提问作者Arnav Mehta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:30:56