非Root Ubuntu用户用Python Tesseract生成可搜索PDF遇FileNotFoundError
解决Ubuntu非Root用户调用Tesseract生成PDF时的FileNotFoundError问题
问题描述
在Ubuntu系统下以非Root用户身份,使用Python调用Tesseract的image_to_pdf_or_hocr生成可搜索PDF时出现FileNotFoundError,但调用image_to_string功能正常。报错详情如下:
FileNotFoundError Traceback (most recent call last) <ipython-input-8-fd512bf1bdc4> in <module> ----> 1 pdf = pt.image_to_pdf_or_hocr(image,lang='hin',extension='pdf') 2 with open('test.pdf', 'w+b') as f: 3 f.write(pdf) ~/hindi_machine_readable/hindi_ocr/lib/python3.6/site-packages/pytesseract/pytesseract.py in image_to_pdf_or_hocr(image, lang, config, nice, extension, timeout) 434 args = [image, extension, lang, config, nice, timeout, True] 435 --> 436 return run_and_get_output(*args) 437 438 ~/hindi_machine_readable/hindi_ocr/lib/python3.6/site-packages/pytesseract/pytesseract.py in run_and_get_output(image, extension, lang, config, nice, timeout, return_bytes) 284 run_tesseract(**kwargs) 285 filename = kwargs['output_filename_base'] + extsep + extension --> 286 with open(filename, 'rb') as output_file: 287 if return_bytes: 288 return output_file.read() FileNotFoundError: [Errno 2] No such file or directory: '/tmp/tess_5il97yg3.pdf'
可能原因及解决方案
1. 缺少PDF生成依赖
image_to_string仅需基础OCR功能,但生成PDF需要Tesseract依赖libtesseract-dev和poppler-utils(负责PDF渲染),非Root用户可能仅安装了基础Tesseract包。
- 有sudo权限的话执行:
sudo apt update && sudo apt install libtesseract-dev poppler-utils - 无sudo权限则联系管理员安装,或从源码编译Tesseract时启用PDF支持。
2. /tmp目录权限异常
报错中的临时文件存储在/tmp,若该目录权限限制了当前用户的读写,会导致文件无法生成。
- 检查目录权限:
正常权限应为ls -ld /tmpdrwxrwxrwt,若不符联系管理员修复。 - 手动指定自定义临时目录(需确保目录有读写权限):
import pytesseract as pt from PIL import Image import os image = Image.open('your_image.jpg') # 创建自定义临时目录 temp_dir = os.path.expanduser('~/my_tess_temp') os.makedirs(temp_dir, exist_ok=True) # 指定临时目录调用接口 pdf = pt.image_to_pdf_or_hocr( image, lang='hin', extension='pdf', config=f'--tmpdir {temp_dir}' ) with open('test.pdf', 'w+b') as f: f.write(pdf)
3. pytesseract版本存在bug
旧版本pytesseract处理临时文件时可能存在逻辑问题,导致PDF临时文件未正确生成或被提前清理。
- 升级到最新版本:
pip install --upgrade pytesseract
4. 印地语语言包不完整
若指定的hin语言包安装不完整,可能导致Tesseract无法生成PDF格式输出。
- 检查已安装语言包:
tesseract --list-langs - 未找到
hin则安装语言包:
无sudo权限可手动下载语言包,放到Tesseract的sudo apt install tesseract-ocr-hintessdata目录(通过tesseract --version查看安装路径)。
内容的提问来源于stack exchange,提问作者PyroSama
相关产品推荐
相关产品推荐

