You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用pytesseract读取PNG图片无法获得正确OCR文本如何解决

问题解决方法
  • 添加图像预处理逻辑
    原始图片直接传入Tesseract识别准确率极低,需要先做灰度化、二值化、降噪处理,过滤干扰像素,提升字符辨识度。
  • 配置Tesseract专用识别参数
    你的识别目标是大写字母+数字的连续串,不需要识别其他特殊字符、多语言内容,可通过限制字符白名单、指定页面分割模式(PSM)大幅提升准确率:
    • 用--psm 8指定识别模式为「识别单个统一的单词」,适配当前短字符串场景
    • 用tessedit_char_whitelist参数限制仅识别大写英文字母和数字,排除乱码字符的输出可能
  • 确认Tesseract运行环境正常
    确保本地安装的Tesseract已自带英文识别包,path_to_tesseract填写的路径无拼写错误、可正常指向Tesseract执行文件。

修改后的完整可运行代码示例:

from PIL import Image, ImageOps
from pytesseract import pytesseract

path_to_tesseract = r"你的Tesseract-OCR.exe实际路径"
image_path = r"你的png文件实际路径"

# 图像预处理流程
img = Image.open(image_path)
# 转灰度图
img = ImageOps.grayscale(img)
# 二值化处理,阈值可根据你的图片实际情况调整
threshold = 127
img = img.point(lambda p: p > threshold and 255 or 0)

pytesseract.tesseract_cmd = path_to_tesseract
# 配置识别参数
custom_config = r'--psm 8 -c tessedit_char_whitelist=ABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789'
text = pytesseract.image_to_string(img, config=custom_config).strip()

print(text)

如果调整后仍存在识别误差,可微调二值化的threshold数值,或者给图片做适当缩放拉伸,让字符高度保持在20-30像素的Tesseract最佳识别范围即可。

内容的提问来源于stack exchange,提问作者Andre_k

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 14:45:02