使用pytesseract提取图像文本返回未知字符,求解决办法
解决pytesseract提取中文图像文本返回乱码的问题
核心问题
你的图像文本清晰,但pytesseract返回乱码,根源是默认未启用中文识别,在线工具内置了多语言支持和适配的预处理逻辑,所以能正常识别。
分步解决方案
1. 强制指定中文语言参数
pytesseract默认仅识别英文,必须添加lang='chi_sim'参数(需提前安装中文语言包):
import pytesseract as pt from PIL import Image # 加载图像 img = Image.open('frame-1ROI_2.png') # 指定中文简体语言进行识别 extracted_text = pt.image_to_string(img, lang='chi_sim') print(extracted_text) print(type(extracted_text))
2. 针对性图像预处理(增强识别率)
如果指定语言后仍有异常,用灰度+二值化强化文本对比度:
import pytesseract as pt from PIL import Image, ImageFilter img = Image.open('frame-1ROI_2.png') # 转灰度图 img = img.convert('L') # 二值化:过滤浅灰色背景,保留黑色文本 threshold = 200 img = img.point(lambda x: 0 if x < threshold else 255, '1') # 轻微锐化提升文本边缘清晰度 img = img.filter(ImageFilter.SHARPEN) extracted_text = pt.image_to_string(img, lang='chi_sim') print(extracted_text)
3. 安装中文语言包
若运行时提示语言包不存在:
- Windows:下载
chi_sim.traineddata放入Tesseract安装目录的tessdata文件夹 - macOS:执行
brew install tesseract-lang - Ubuntu/Debian:执行
sudo apt install tesseract-ocr-chi-sim
内容的提问来源于stack exchange,提问作者JAMSHAID
相关产品推荐
相关产品推荐

