You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pytesseract提取图像文本返回未知字符,求解决办法

解决pytesseract提取中文图像文本返回乱码的问题

核心问题

你的图像文本清晰,但pytesseract返回乱码,根源是默认未启用中文识别,在线工具内置了多语言支持和适配的预处理逻辑,所以能正常识别。

分步解决方案

1. 强制指定中文语言参数

pytesseract默认仅识别英文,必须添加lang='chi_sim'参数(需提前安装中文语言包):

import pytesseract as pt
from PIL import Image

# 加载图像
img = Image.open('frame-1ROI_2.png')
# 指定中文简体语言进行识别
extracted_text = pt.image_to_string(img, lang='chi_sim')
print(extracted_text)
print(type(extracted_text))

2. 针对性图像预处理(增强识别率)

如果指定语言后仍有异常,用灰度+二值化强化文本对比度:

import pytesseract as pt
from PIL import Image, ImageFilter

img = Image.open('frame-1ROI_2.png')
# 转灰度图
img = img.convert('L')
# 二值化:过滤浅灰色背景,保留黑色文本
threshold = 200
img = img.point(lambda x: 0 if x < threshold else 255, '1')
# 轻微锐化提升文本边缘清晰度
img = img.filter(ImageFilter.SHARPEN)

extracted_text = pt.image_to_string(img, lang='chi_sim')
print(extracted_text)

3. 安装中文语言包

若运行时提示语言包不存在:

  • Windows:下载chi_sim.traineddata放入Tesseract安装目录的tessdata文件夹
  • macOS:执行brew install tesseract-lang
  • Ubuntu/Debian:执行sudo apt install tesseract-ocr-chi-sim

内容的提问来源于stack exchange,提问作者JAMSHAID

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 15:24:06