使用python-tesseract的PSM 7读取含重叠签名的图片返回空结果求助
解决Tesseract识别带重叠签名姓名的问题
首先修正代码配置遗漏问题
你已经定义了custom_oem_psm_config参数,但没有传递给image_to_string方法,导致PSM 7模式根本没生效,这是返回空结果的核心原因之一。修正后的代码如下:
pytesseract.pytesseract.tesseract_cmd = r'C:\Users\hoffmann\AppData\Local\anaconda3\Library\bin\tesseract.exe' custom_oem_psm_config = r'--oem 3 --psm 7' image = Image.open('G:/PraktikantInnen/PRAKTIKANTINNEN/Hoffmann/Read_pdf_Stella/face.png') # 传入配置参数启用PSM 7模式 text = pytesseract.image_to_string(image, config=custom_oem_psm_config) print(text)
图像预处理优化(若修正配置后仍无结果)
签名重叠会干扰字符特征,可通过预处理弱化干扰:
- 二值化+降噪:将图像转为黑白并过滤淡色签名,强化姓名字符:
from PIL import Image, ImageFilter # 转灰度图 gray_img = image.convert('L') # 二值化(调整阈值过滤淡色签名,可根据实际图像亮度修改阈值) threshold = 180 binary_img = gray_img.point(lambda x: 0 if x < threshold else 255, '1') # 轻微降噪消除杂点 cleaned_img = binary_img.filter(ImageFilter.MedianFilter(size=3)) # 用处理后的图像识别 text = pytesseract.image_to_string(cleaned_img, config=custom_oem_psm_config) - 精准裁剪区域:如果能确定姓名的大致位置,直接裁剪掉签名覆盖的部分,只保留姓名区域再识别,大幅降低干扰。
补充语言包支持
如果识别的是德语姓名(如示例中的Max Müller),需确保Tesseract已安装德语语言包,并在识别时指定语言:
text = pytesseract.image_to_string(cleaned_img, lang='deu', config=custom_oem_psm_config)
内容的提问来源于stack exchange,提问作者Jojo
相关产品推荐
相关产品推荐

