You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:pytesseract识别本地图片文本返回错误值问题

图片文本识别不准确的解决求助

我尝试识别本地存储的多张图片中的文本,但返回结果始终不准确:

  • 预期结果e4cbd,实际返回‘edchd
  • 预期结果4x22m,实际返回AL

最初使用以下代码时,甚至没有任何输出:

import pytesseract
from PIL import Image
image = Image.open('abc.png')
text = pytesseract.image_to_string(image)
print(text)

修改为text = pytesseract.image_to_string(img, lang='eng', config='--psm 6')后,虽能得到结果,但识别准确率依然不达标。测试图片对应的预期文本分别为:

  • e4cbd
  • 4x22m
  • grmmc
  • 2p6y8
  • faxy4

解决建议

这类短字符(类似验证码)的识别问题,核心在于图像预处理和Tesseract参数优化,以下是具体方案:

1. 图像预处理

通过灰度转换、二值化、颜色反转等操作,强化字符与背景的对比度,帮助Tesseract更精准识别:

import pytesseract
from PIL import Image, ImageOps

def preprocess_image(image_path):
    # 打开图片并转换为灰度模式
    img = Image.open(image_path).convert('L')
    # 二值化处理,调整阈值让字符边缘更清晰
    threshold = 127
    img = img.point(lambda pixel: 255 if pixel > threshold else 0)
    # 反转颜色(若字符为浅色、背景为深色时适用)
    img = ImageOps.invert(img)
    return img

2. 优化Tesseract识别参数

针对短字符场景,调整参数限定识别范围和模式:

# 预处理图片
processed_img = preprocess_image('abc.png')
# 配置参数:单个字符块识别 + 限定字符范围 + 默认引擎
text = pytesseract.image_to_string(
    processed_img,
    lang='eng',
    config='--psm 10 --oem 3 -c tessedit_char_whitelist=abcdefghijklmnopqrstuvwxyz0123456789'
)
# 去除多余空白后输出
print(text.strip())

参数说明:

  • --psm 10:指定识别单个字符(适合短验证码类场景)
  • --oem 3:使用默认OCR引擎组合模式
  • tessedit_char_whitelist:仅识别小写字母和数字,排除无关字符干扰

3. 验证预处理效果

可以将处理后的图片保存,确认字符是否清晰可见:

processed_img.save('processed_image.png')

内容的提问来源于stack exchange,提问作者ritika kapoor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 05:42:53