Tesseract-OCR LSTM训练:多文件列表失效单文件正常问题求助
Tesseract手写识别微调训练故障排查
问题描述
我在微调Tesseract手写识别模型时遇到了问题:已经准备好字符图像和对应box文件,生成了.lstmf文件,也用Python脚本生成了lstm_train.txt和lstm_test.txt,但用这两个列表文件启动训练就失败,只有当列表里只留单个.lstmf文件路径时才能正常启动。另外,所有.lstmf文件本身是没问题的——我写了个逐文件训练、从上次checkpoint续训的脚本,能正常处理所有文件。现在不确定是lstm_train.txt的生成环节有问题,还是lstmtraining工具本身只支持单个.lstmf文件作为输入?
生成列表文件的Python代码
import os import random input_dir = "test" train_file = "lstm_train.txt" test_file = "lstm_test.txt" # 列出所有.lstmf文件 all_files = [f for f in os.listdir(input_dir) if f.endswith(".lstmf")] random.shuffle(all_files) # 随机打乱 # 训练集比例(80%) train_split = 0.8 train_count = int(len(all_files) * train_split) train_files = all_files[:train_count] test_files = all_files[train_count:] # 写入训练和测试文件,使用相对路径 with open(train_file, "w", encoding="utf-8") as f_train, \ open(test_file, "w", encoding="utf-8") as f_test: for f in train_files: relative_path = os.path.join(input_dir, f) f_train.write(relative_path+"\n") for f in test_files: relative_path = os.path.join(input_dir, f) f_test.write(relative_path+"\n") print(f"[OK] 文件 '{train_file}' 和 '{test_file}' 已生成,使用相对路径。")
排查与解决建议
- 检查路径格式:确保lstm_train.txt里的路径没有多余空格、换行符,或包含特殊字符(如中文、空格)。路径含空格时需用引号包裹,或替换为无空格路径。
- 验证文件编码:确认列表文件是UTF-8无BOM编码,避免编辑器自动添加的BOM头导致解析失败。可通过记事本选择UTF-8编码重新保存。
- 清理列表文件格式:确保每个.lstmf路径单独占一行,无空行或重复行。可手动检查,或用命令
cat lstm_train.txt | grep -v "^$"过滤空行后重试。 - 升级Tesseract版本:旧版本Tesseract处理多文件列表可能存在bug,建议升级到5.3.0及以上的稳定版本。
- 改用绝对路径:将脚本中的相对路径替换为绝对路径,避免工作目录问题导致Tesseract找不到文件。修改后的核心代码片段如下:
# 替换原路径写入部分 for f in train_files: abs_path = os.path.abspath(os.path.join(input_dir, f)) f_train.write(f"{abs_path}\n") for f in test_files: abs_path = os.path.abspath(os.path.join(input_dir, f)) f_test.write(f"{abs_path}\n")
内容的提问来源于stack exchange,提问作者TestING
相关产品推荐
相关产品推荐

