You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用python-tesseract的PSM 7读取含重叠签名的图片返回空结果求助

解决Tesseract识别带重叠签名姓名的问题

首先修正代码配置遗漏问题

你已经定义了custom_oem_psm_config参数,但没有传递给image_to_string方法,导致PSM 7模式根本没生效,这是返回空结果的核心原因之一。修正后的代码如下:

pytesseract.pytesseract.tesseract_cmd = r'C:\Users\hoffmann\AppData\Local\anaconda3\Library\bin\tesseract.exe'
custom_oem_psm_config = r'--oem 3 --psm 7'
image = Image.open('G:/PraktikantInnen/PRAKTIKANTINNEN/Hoffmann/Read_pdf_Stella/face.png')
# 传入配置参数启用PSM 7模式
text = pytesseract.image_to_string(image, config=custom_oem_psm_config)
print(text)

图像预处理优化(若修正配置后仍无结果)

签名重叠会干扰字符特征,可通过预处理弱化干扰:

  • 二值化+降噪:将图像转为黑白并过滤淡色签名,强化姓名字符:
    from PIL import Image, ImageFilter
    
    # 转灰度图
    gray_img = image.convert('L')
    # 二值化(调整阈值过滤淡色签名,可根据实际图像亮度修改阈值)
    threshold = 180
    binary_img = gray_img.point(lambda x: 0 if x < threshold else 255, '1')
    # 轻微降噪消除杂点
    cleaned_img = binary_img.filter(ImageFilter.MedianFilter(size=3))
    # 用处理后的图像识别
    text = pytesseract.image_to_string(cleaned_img, config=custom_oem_psm_config)
    
  • 精准裁剪区域:如果能确定姓名的大致位置,直接裁剪掉签名覆盖的部分,只保留姓名区域再识别,大幅降低干扰。

补充语言包支持

如果识别的是德语姓名(如示例中的Max Müller),需确保Tesseract已安装德语语言包,并在识别时指定语言:

text = pytesseract.image_to_string(cleaned_img, lang='deu', config=custom_oem_psm_config)

内容的提问来源于stack exchange,提问作者Jojo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 14:05:13