使用pytesseract图片转文本遇TesseractNotFoundError错误求助
问题描述
我运行以下Python脚本尝试用OCR识别图片文本:
from PIL import Image print(pytesseract.image_to_string(Image.open('/Users/Sanya/Downloads/SimpleText.png')))
但出现了如下错误:
Traceback (most recent call last): File "/Library/Frameworks/Python.framework/Versions/3.8/lib/python3.8/site-packages/pytesseract/pytesseract.py", line 250, in run_tesseract proc = subprocess.Popen(cmd_args, **subprocess_args()) File "/Library/Frameworks/Python.framework/Versions/3.8/lib/python3.8/subprocess.py", line 854, in __init__ self._execute_child(args, executable, preexec_fn, close_fds, File "/Library/Frameworks/Python.framework/Versions/3.8/lib/python3.8/subprocess.py", line 1702, in _execute_child raise child_exception_type(errno_num, err_msg, err_filename) FileNotFoundError: [Errno 2] No such file or directory: 'tesseract' During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/Users/sanya/Documents/PythonFile/ImgToText.py", line 8, in <module> pytesseract.image_to_string(img, config=custom_config) File "/Library/Frameworks/Python.framework/Versions/3.8/lib/python3.8/site-packages/pytesseract/pytesseract.py", line 370, in image_to_string return { File "/Library/Frameworks/Python.framework/Versions/3.8/lib/python3.8/site-packages/pytesseract/pytesseract.py", line 373, in <lambda> Output.STRING: lambda: run_and_get_output(*args), File "/Library/Frameworks/Python.framework/Versions/3.8/lib/python3.8/site-packages/pytesseract/pytesseract.py", line 282, in run_and_get_output run_tesseract(**kwargs) File "/Library/Frameworks/Python.framework/Versions/3.8/lib/python3.8/site-packages/pytesseract/pytesseract.py", line 254, in run_tesseract raise TesseractNotFoundError() pytesseract.pytesseract.TesseractNotFoundError: tesseract is not installed or it's not in your PATH. See README file for more information.
尝试了一些Tesseract相关操作后还是出现相同错误,求解决办法。
解决方案
这个错误的核心原因很明确:你只安装了Python的pytesseract封装库,但没有安装底层的Tesseract OCR引擎本身,或者引擎已经安装但没有被添加到系统的环境变量PATH中,导致Python找不到它。
下面分步骤解决,针对你的Mac环境(从路径能看出来)优先说明:
1. 先确认Tesseract是否已安装
打开终端,输入以下命令:
which tesseract
如果终端没有任何输出,说明你还没安装Tesseract;如果输出了类似/usr/local/bin/tesseract的路径,说明已安装,但可能PATH配置有问题,直接跳到步骤3。
2. 安装Tesseract OCR引擎
对于Mac,最方便的方式是用Homebrew包管理器:
- 如果还没装Homebrew,先在终端执行:
/usr/bin/ruby -e "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/master/install)" - 然后安装Tesseract:
brew install tesseract
如果不想用Homebrew,也可以从Tesseract官方仓库下载对应Mac的DMG安装包手动安装,但Homebrew的方式更省心,还能自动处理依赖。
3. 验证安装是否成功
安装完成后,在终端输入:
tesseract --version
如果能输出Tesseract的版本信息(比如tesseract 5.3.3),说明安装成功。
4. 若仍报错,手动指定Tesseract路径
有时候即使安装了,Python可能还是找不到Tesseract的位置,这时候可以在你的Python代码里手动指定它的路径:
from PIL import Image import pytesseract # 替换成你的Tesseract实际路径,Mac下默认是/usr/local/bin/tesseract pytesseract.pytesseract.tesseract_cmd = '/usr/local/bin/tesseract' print(pytesseract.image_to_string(Image.open('/Users/Sanya/Downloads/SimpleText.png')))
其他平台的补充方案(如果需要)
- Windows:从Tesseract官方渠道下载安装包,安装时一定要勾选"Add to PATH"选项;如果没勾选,手动把安装目录(比如
C:\Program Files\Tesseract-OCR)添加到系统环境变量PATH中。 - Linux(Ubuntu/Debian):执行
sudo apt update && sudo apt install tesseract-ocr;CentOS/RHEL则用sudo yum install tesseract。
按照这些步骤操作后,你的OCR脚本应该就能正常运行了。
内容的提问来源于stack exchange,提问作者Strange
相关产品推荐
相关产品推荐

