You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化图像预处理效果,提升Tesseract OCR的文本识别完整性

问题描述

我有如下待识别图像:
待识别原始图像

为了尽可能达到最佳的Tesseract识别效果,我采用了如下流程处理图像:

sharpened = unsharp_mask(img, amount=1.5)
cv2.imwrite(TEMP_FOLDER + 'sharpened.png', sharpened)

thr = cv2.threshold(sharpened, 220, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]
# im = cv2.resize(thr, None, fx=2, fy=2, interpolation=cv2.INTER_AREA)
os.makedirs(TEMP_FOLDER, exist_ok=True)
cv2.imwrite(TEMP_FOLDER + 'inverted.png', thr)
inverted = cv2.imread(TEMP_FOLDER + 'inverted.png')

filtered_inverted = remove_black_boundaries(inverted)
filtered_inverted = cv2.resize(filtered_inverted, None, fx=2, fy=2, interpolation=cv2.INTER_LINEAR)
# kernel = np.ones((2, 2), np.uint8)
# filtered_inverted = cv2.dilate(filtered_inverted, kernel)
cv2.imwrite(TEMP_FOLDER + 'filtered.png', filtered_inverted)

median = cv2.medianBlur(filtered_inverted, 5)
# median = cv2.cvtColor(median, cv2.COLOR_RGB2GRAY)
# median = cv2.threshold(median, 127, 255, cv2.THRESH_BINARY)[1]
cv2.imwrite(TEMP_FOLDER + 'median.png', median)

其中unsharp_mask函数的定义如下:

def unsharp_mask(image: np.ndarray, kernel_size: Tuple[int] = (5, 5),
             sigma: float = 1.0, amount: float = 1.0, threshold: float = 0) -> np.ndarray:
    """Return a sharpened version of the image, using an unsharp mask."""
    blurred = cv2.GaussianBlur(image, kernel_size, sigma)
    sharpened = float(amount + 1) * image - float(amount) * blurred
    sharpened = np.maximum(sharpened, np.zeros(sharpened.shape))
    sharpened = np.minimum(sharpened, 255 * np.ones(sharpened.shape))
    sharpened = sharpened.round().astype(np.uint8)
    if threshold > 0:
        low_contrast_mask = np.absolute(image - blurred) < threshold
        np.copyto(sharpened, image, where=low_contrast_mask)
    return sharpened

remove_black_boundaries函数(本次场景中图像无黑色边界,该函数未生效)定义如下:

def remove_black_boundaries(img: np.ndarray) -> np.ndarray:
    hh, ww = img.shape[:2]
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    _, thresh = cv2.threshold(gray, 1, 255, cv2.THRESH_BINARY)
    thresh = cv2.erode(thresh, np.ones((3, 3), np.uint8))

    contours = cv2.findContours(thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    contours = contours[0] if len(contours) == 2 else contours[1]
    cnt = max(contours, key=cv2.contourArea)
    # draw white contour on black background as mask
    mask = np.zeros((hh, ww), dtype=np.uint8)
    cv2.drawContours(mask, [cnt], 0, (255, 255, 255), cv2.FILLED)

    # invert mask so shapes are white on black background
    mask_inv = 255 - mask

    # create new (white) background
    bckgnd = np.full_like(img, (255, 255, 255))

    # apply mask to image
    image_masked = cv2.bitwise_and(img, img, mask=mask)

    # apply inverse mask to background
    bckgnd_masked = cv2.bitwise_and(bckgnd, bckgnd, mask=mask_inv)

    # add together
    result = cv2.add(image_masked, bckgnd_masked)
    return result

经过上述处理,各步骤输出图像如下:

  • 锐化图像:锐化图像
  • 反色并过滤后的图像:反色过滤后图像
  • 最终输入Tesseract的图像:最终输入图像

目前Tesseract仅能识别出Conteggio: 2900,缺失第一行文本,调整图像尺寸后结果依旧,如何优化输入Tesseract的图像质量以实现完整识别?


解决方案

识别缺失第一行的核心原因是现有处理流程过度过滤了低对比度的细笔画字符,可按以下步骤优化:

  • 替换全局二值化为自适应阈值:全局阈值220会把第一行浅灰色字符直接判定为背景过滤,改用自适应局部阈值适配不同区域的亮度差异,代码参考:
    gray = cv2.cvtColor(sharpened, cv2.COLOR_BGR2GRAY)
    thr = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    
  • 移除大核中值模糊:5x5的中值模糊核会把第一行的细笔画直接磨断甚至完全消除,最多保留3x3的高斯模糊做轻微降噪即可,不需要做中值模糊处理。
  • 调整锐化参数:将unsharp_mask的amount参数提升到2.0,kernel_size改为(3,3),针对性增强细笔画的对比度,同时避免过度锐化产生噪点。
  • 增加小角度倾斜校正:图像存在轻微的逆时针倾斜,可通过minAreaRect检测文本轮廓做±5度范围内的旋转校正,匹配Tesseract的水平文本识别偏好。
  • 优化Tesseract调用参数:识别时添加--psm 6参数(假设输入为单块统一文本),如果目标文本为意大利语可额外添加-l ita指定语言包,进一步提升识别准确率。

内容的提问来源于stack exchange,提问作者marco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 23:27:05