Tesseract lstmtraining--continue_from报错:希伯来语手写字体训练失败
解决Tesseract LSTM训练时"Failed to read continue from: output/.tr"错误
错误根源分析
你遇到的错误有两个核心原因:
- 脚本中
lstmtraining命令被注释(前面的#导致命令未执行),后续参数被错误解析,引发路径读取异常; - 错误地将旧格式训练文件
.tr传入--continue_from参数,该参数要求的是LSTM预训练模型文件(.lstm或.checkpoint格式),而非.tr文件。
修正后的完整训练流程
针对你的希伯来语手写字体训练场景,按以下步骤执行:
1. 确认前期文件生成正常
检查text2image和tesseract box.train步骤已成功生成output/2.tif、output/2.box、output/2.tr文件,无报错。
2. 提取预训练LSTM模型
从你已有的best版本heb.traineddata中提取LSTM核心模型:
combine_tessdata -e tessdata/heb.traineddata output/heb.lstm
3. 生成LSTM训练专用的.lstmf文件
将生成的tif转换为LSTM训练所需的格式:
tesseract output/2.tif output/2 -l heb lstm.train
执行后会得到output/2.lstmf文件。
4. 执行LSTM微调训练
lstmtraining ^ --continue_from output/heb.lstm ^ --traineddata tessdata/heb.traineddata ^ --train_listfile output/2.lstmf ^ --model_output output/heb_handwriting ^ --max_iterations 1000
(Windows环境用^作为换行符,若用WSL则替换为\)
5. 停止训练并生成最终模型
训练迭代完成后,生成可用的.traineddata文件:
lstmtraining ^ --stop_training ^ --continue_from output/heb_handwriting_checkpoint ^ --traineddata tessdata/heb.traineddata ^ --model_output output/heb_handwriting.traineddata
额外注意事项
- 检查Windows路径是否存在空格或特殊字符,必要时使用绝对路径;
- 确保
langdata_lstm/heb目录下的语言数据文件完整; - 若
max_pages=2仅为POC测试,后续可增大页数提升模型效果。
内容的提问来源于stack exchange,提问作者roni
相关产品推荐
相关产品推荐

