在Heroku部署含ocrmypdf的Node.js服务器时遇Tesseract断言错误
我在Heroku上部署一个使用OCRmyPDF的Node.js Web服务器,当前使用的Heroku buildpack如下:
- heroku/python
- https://github.com/heroku/heroku-buildpack-apt
- https://github.com/pathwaysmedical/heroku-buildpack-tesseract
- heroku/nodejs
Tesseract运行正常,但OCRmyPDF无法工作。我的requirements.txt文件仅配置了ocrmypdf==13。
执行命令ocrmypdf l.pdf l2.pdf后出现以下错误日志:
Opened a file Opened a file Opened a file Opened a file Opened a file Opened a file Opened a file Scanning contents: 100%|███████████████████████████████████████████████████████████████████████████ 20/20 [00:00<00:00, 381.51page/s] Start processing 8 pages concurrently Opened a file 1 [tesseract] read_params_file: Can't open pdf 1 [tesseract] read_params_file: Can't open txt 1 [tesseract] Warning: Invalid resolution 25 dpi. Using 70 instead. 1 [tesseract] Estimating resolution as 418 3 [tesseract] read_params_file: Can't open pdf 3 [tesseract] read_params_file: Can't open txt 3 [tesseract] Warning: Invalid resolution 25 dpi. Using 70 instead. 3 [tesseract] Estimating resolution as 215 2 [tesseract] read_params_file: Can't open pdf 2 [tesseract] read_params_file: Can't open txt 2 [tesseract] Warning: Invalid resolution 25 dpi. Using 70 instead. 2 [tesseract] Estimating resolution as 236 4 [tesseract] read_params_file: Can't open pdf 4 [tesseract] read_params_file: Can't open txt 4 [tesseract] Warning: Invalid resolution 25 dpi. Using 70 instead. 4 [tesseract] Estimating resolution as 224 4 [tesseract] contains_unichar_id(unichar_id):Error:Assert failed:in file ../../src/ccutil/unicharset.h, line 509 6 [tesseract] read_params_file: Can't open pdf 6 [tesseract] read_params_file: Can't open txt 6 [tesseract] Warning: Invalid resolution 25 dpi. Using 70 instead. 6 [tesseract] Estimating resolution as 199 OCR: 20%|█████████████████████████████████████▏ | 4.0/20.0 [00:05<00:20, 1.30s/page] SubprocessOutputError
1. 补充Apt依赖包
OCRmyPDF依赖Poppler、Tesseract开发库等底层工具,在项目根目录创建Aptfile文件,添加以下内容:
poppler-utils libtesseract-dev tesseract-ocr-eng
如果需要其他语言的OCR支持,添加对应语言包,比如tesseract-ocr-chi-sim(简体中文)。
2. 更换Tesseract Buildpack
你当前使用的第三方buildpack可能版本老旧或缺少必要组件,替换为Heroku官方维护的Tesseract buildpack:
heroku buildpacks:remove https://github.com/pathwaysmedical/heroku-buildpack-tesseract heroku buildpacks:add https://github.com/heroku/heroku-buildpack-tesseract
该buildpack会安装稳定版Tesseract,并确保语言包路径配置正确。
3. 调整OCRmyPDF执行参数
日志中的分辨率警告可能触发Tesseract的断言错误,执行OCR时强制指定dpi,并禁用并发排查冲突:
ocrmypdf --dpi 300 --jobs 1 l.pdf l2.pdf
--dpi 300解决分辨率异常问题,--jobs 1避免多进程并发导致的资源或依赖冲突。
4. 验证依赖版本兼容性
OCRmyPDF 13要求Tesseract版本≥4.1.0,在Heroku控制台执行以下命令确认版本:
tesseract --version
如果版本过低,确认buildpack安装的是正确版本,或通过环境变量指定Tesseract版本(部分buildpack支持此配置)。
5. 修正Buildpack顺序
确保buildpack按依赖优先级排序:Apt包先安装,再依次是Python、Tesseract、Node.js:
heroku buildpacks:set --index 1 heroku/python heroku buildpacks:set --index 2 https://github.com/heroku/heroku-buildpack-apt heroku buildpacks:set --index 3 https://github.com/heroku/heroku-buildpack-tesseract heroku buildpacks:set --index 4 heroku/nodejs
内容的提问来源于stack exchange,提问作者Vadzim Papkou

