为何pytesseract无法识别目标图片?已尝试多种预处理仍无效
问题解决:pytesseract无法识别指定图片文字
问题说明
pytesseract可正常识别其他图片,但无法识别目标图片。已尝试调整尺寸、转灰度图、黑白反转等预处理操作,均无效果。目标图片特征为浅色文字在深色渐变背景上。
用户原有代码
import numpy as np from pytesseract import pytesseract import cv2 import os path_to_tesseract = r'C:\Program Files\Tesseract-OCR\tesseract.exe' pytesseract.tesseract_cmd = path_to_tesseract img = cv2.imread(r'C:\Users\Owner\Desktop\Coding\PNGs\tugteam project\tugteam2.png') grayImage = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) (thresh, img) = cv2.threshold(grayImage, 127, 255, cv2.THRESH_BINARY) img = cv2.bitwise_not(img) img = cv2.resize(img, (600, 400)) cv2.imshow('asd',img) cv2.waitKey(0) cv2.destroyAllWindows() text = pytesseract.image_to_string(img) print(text)
解决方法及优化代码
1. 替换固定阈值为自适应阈值
目标图片背景是渐变的,固定阈值会导致部分文字丢失。用自适应阈值能根据局部区域调整阈值,完整保留文字轮廓。
2. 添加形态学膨胀操作
填补文字边缘的微小缺口,让文字轮廓更清晰,降低识别难度。
3. 指定Tesseract配置参数
设置--psm 6(假设输入为单一均匀文本块)和--oem 3(使用默认引擎模式),针对性提升识别精准度。
修改后代码:
import numpy as np from pytesseract import pytesseract import cv2 import os path_to_tesseract = r'C:\Program Files\Tesseract-OCR\tesseract.exe' pytesseract.tesseract_cmd = path_to_tesseract # 读取图片 img = cv2.imread(r'C:\Users\Owner\Desktop\Coding\PNGs\tugteam project\tugteam2.png') gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # 自适应阈值处理,替代固定阈值 thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) # 形态学膨胀操作,增强文字轮廓 kernel = np.ones((2,2), np.uint8) thresh = cv2.dilate(thresh, kernel, iterations=1) # 放大图片(可选,提升小文字识别率) thresh = cv2.resize(thresh, None, fx=2, fy=2, interpolation=cv2.INTER_CUBIC) # 显示预处理后的图片 cv2.imshow('Processed Image', thresh) cv2.waitKey(0) cv2.destroyAllWindows() # 设置Tesseract配置参数 custom_config = r'--oem 3 --psm 6' text = pytesseract.image_to_string(thresh, config=custom_config) print("识别结果:") print(text)
额外提示
- 若仍有识别问题,可添加高斯模糊降噪:
gray = cv2.GaussianBlur(gray, (3,3), 0),减少背景噪点干扰。 - 确认Tesseract已安装对应语言包(英文包默认已安装,识别其他语言需额外安装)。
内容的提问来源于stack exchange,提问作者twitterL9
相关产品推荐
相关产品推荐

