You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Tesseract无法识别特定类型图片中数字问题求助

问题排查及解决方案

核心问题原因

  • 像素比较逻辑错误:你将图片转为LA模式(灰度+alpha通道),每个像素为二元组(灰度值, 透明度),但你用该值与RGB三元组阈值(160,160,160)对比,比较逻辑完全不符合预期,直接导致二值化结果异常,生成大量噪点。
  • 对比度增强操作未生效:你调用enh.enhance(1.3)仅做了展示,后续二值化处理使用的仍是未增强的out对象,对比度提升操作没有实际作用。
  • 页分割模式(PSM)配置错误:你使用--psm 10参数,该参数是告诉Tesseract待识别内容为单个字符,若你的待识别图片包含多个连续数字,该配置会直接导致识别逻辑错误。你配置了数字白名单却识别出字母,正是PSM配置错误导致白名单规则未正常生效。可根据数字排布选择PSM参数:单行数字用7,分散的多个数字用6,单个数字保留10即可。
  • 固定阈值适配性差:新的待识别图片大概率存在渐变背景、浅色干扰噪点,固定阈值160无法适配这类图片的二值化需求,会导致数字轮廓被破坏或者背景噪点被识别为字符。

修正方案

版本1:引入OpenCV做自适应阈值处理(识别准确率更高)

需先安装依赖:pip install opencv-python numpy

try:
    from PIL import Image
    from PIL import ImageEnhance
except ImportError:
    import Image
import pytesseract
import cv2
import numpy as np

def binarize_image(img_path):
    # 灰度读入
    img = cv2.imread(img_path, 0)
    # 自适应高斯阈值二值化,适配带干扰的背景
    thresh = cv2.adaptiveThreshold(img, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    # 形态学操作去除细碎噪点
    kernel = np.ones((2, 2), np.uint8)
    cleaned = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=1)
    return Image.fromarray(cleaned)

if __name__ == "__main__":
    input_path = "./in/web_search.jpg"
    processed_img = binarize_image(input_path)
    # 可选二次提升对比度
    enh = ImageEnhance.Contrast(processed_img)
    processed_img = enh.enhance(1.5)
    
    pytesseract.pytesseract.tesseract_cmd = r'/usr/bin/tesseract'
    print("-----------------------")
    # 调整PSM为7,适配单行数字识别场景
    print(pytesseract.image_to_string(processed_img, lang='eng', config='--psm 7 --oem 3 -c tessedit_char_whitelist=1234567890 --tessdata-dir="/usr/share/tesseract-ocr/4.00/tessdata/"'))
    print("-----------------------")
    # 保存处理后的图片可自行校验效果
    processed_img.save("./out/web_search_processed.jpg")

版本2:纯PIL实现(无额外依赖)

try:
    from PIL import Image
    from PIL import ImageEnhance
except ImportError:
    import Image
import pytesseract

if __name__ == "__main__":
    input_path = "./in/web_search.jpg"
    # 直接转灰度L模式,去掉无用的alpha通道
    img = Image.open(input_path).convert("L")
    # 提升亮度
    img = img.point(lambda x: x * 1.4)
    # 对比度增强操作结果直接用于后续处理
    enh = ImageEnhance.Contrast(img)
    img = enh.enhance(1.4)
    
    # 正确的单通道阈值判断逻辑
    threshold = 160
    pixels = img.getdata()
    new_pixels = [0 if p < threshold else 255 for p in pixels]
    new_img = Image.new("L", img.size)
    new_img.putdata(new_pixels)
    
    pytesseract.pytesseract.tesseract_cmd = r'/usr/bin/tesseract'
    print("-----------------------")
    # 调整PSM为7,适配单行数字识别场景
    print(pytesseract.image_to_string(new_img, lang='eng', config='--psm 7 --oem 3 -c tessedit_char_whitelist=1234567890 --tessdata-dir="/usr/share/tesseract-ocr/4.00/tessdata/"'))
    print("-----------------------")
    new_img.save("./out/web_search_processed.jpg")

内容的提问来源于stack exchange,提问作者Dev Dev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 15:45:04