微调PaddleOCR识别模型时遭遇预训练参数不匹配警告及递归错误的求助
微调PaddleOCR识别模型时遭遇预训练参数不匹配警告及递归错误的求助
我最近一直在尝试微调PaddleOCR的识别模型,用来处理OCR-A字体的字符——默认模型对这类字体识别不准,比如会把1认成L,6认成b。我查了好多文章和官方文档,但还是遇到一堆问题,而且感觉文档写得有点模糊。
我的操作流程
- 用PPOCRLabel标注图片,生成了
Label.txt标注文件; - 基于
rec_mv3_none_bilstm_ctc配置文件修改适配我的数据集; - 执行的训练命令:
!python3 /content/PaddleOCR/tools/train.py -c /content/config.yml -o Global.pretrained_model=/content/pretrain_models_rec/en_PP-OCRv3_rec_train/best_accuracy
遇到的问题
1. 预训练参数不匹配警告
命令执行后先弹出了一串参数不匹配的警告,虽然最后提示预训练模型加载成功,但还是担心有问题:
[2024/12/09 03:43:11] ppocr WARNING: The pretrained params head.sar_head.decoder.rnn_decoder.1.cell.weight_ih not in model [2024/12/09 03:43:11] ppocr WARNING: The pretrained params head.sar_head.decoder.rnn_decoder.1.cell.weight_hh not in model [2024/12/09 03:43:11] ppocr WARNING: The pretrained params head.sar_head.decoder.rnn_decoder.1.cell.bias_ih not in model [2024/12/09 03:43:11] ppocr WARNING: The pretrained params head.sar_head.decoder.rnn_decoder.1.cell.bias_hh not in model [2024/12/09 03:43:11] ppocr WARNING: The pretrained params head.sar_head.decoder.embedding.weight not in model [2024/12/09 03:43:11] ppocr WARNING: The pretrained params head.sar_head.decoder.prediction.weight not in model [2024/12/09 03:43:11] ppocr WARNING: The pretrained params head.sar_head.decoder.prediction.bias not in model [2024/12/09 03:43:11] ppocr INFO: load pretrain successful from /content/pretrain_models_rec/en_PP-OCRv3_rec_train/best_accuracy
2. 致命的递归错误
警告之后,程序直接抛出了递归错误,训练中断,完整报错栈如下:
Exception in thread Thread-1 (_thread_loop): Traceback (most recent call last): File "/usr/lib/python3.10/threading.py", line 1016, in _bootstrap_inner self.run() File "/usr/lib/python3.10/threading.py", line 953, in run self._target(*self._args, **self._kwargs) File "/usr/local/lib/python3.10/dist-packages/paddle/io/dataloader/dataloader_iter.py", line 603, in _thread_loop batch = self._get_data() File "/usr/local/lib/python3.10/dist-packages/paddle/io/dataloader/dataloader_iter.py", line 752, in _get_data batch.reraise() File "/usr/local/lib/python3.10/dist-packages/paddle/io/dataloader/worker.py", line 187, in reraise raise self.exc_type(msg) RecursionError: DataLoader worker(0) caught RecursionError with message: Traceback (most recent call last): File "/usr/local/lib/python3.10/dist-packages/paddle/io/dataloader/worker.py", line 372, in _worker_loop batch = fetcher.fetch(indices) File "/usr/local/lib/python3.10/dist-packages/paddle/io/dataloader/fetcher.py", line 77, in fetch data.append(self.dataset[idx]) File "/content/PaddleOCR/ppocr/data/simple_dataset.py", line 163, in __getitem__ return self.__getitem__(rnd_idx) File "/content/PaddleOCR/ppocr/data/simple_dataset.py", line 163, in __getitem__ return self.__getitem__(rnd_idx) File "/content/PaddleOCR/ppocr/data/simple_dataset.py", line 163, in __getitem__ return self.__getitem__(rnd_idx) [Previous line repeated 7 more times] File "/content/PaddleOCR/ppocr/data/simple_dataset.py", line 161, in __getitem__ raise RecursionError("Maximum recursion depth exceeded in __getitem__") RecursionError: Maximum recursion depth exceeded in __getitem__ Traceback (most recent call last): File "/content/PaddleOCR/tools/train.py", line 269, in <module> main(config, device, logger, vdl_writer, seed) File "/content/PaddleOCR/tools/train.py", line 222, in main program.train( File "/content/PaddleOCR/tools/program.py", line 312, in train for idx, batch in enumerate(train_dataloader): File "/usr/local/lib/python3.10/dist-packages/paddle/io/dataloader/dataloader_iter.py", line 826, in __next__ self._reader.read_next_list()[0] SystemError: (Fatal) Blocking queue is killed because the data reader raises an exception. [Hint: Expected killed_ != true, but received killed_:1 == true:1.] (at /paddle/paddle/fluid/operators/reader/blocking_queue.h:175)
已尝试的排查动作
我已经反复检查过配置文件里的图片路径、标注文件路径,确认路径是正确的,但问题依然存在。
想请教的问题
- 这些预训练参数不匹配的警告是什么原因导致的?会不会影响后续的训练效果?
- 这个递归错误到底怎么解决?看起来是Dataset的
__getitem__方法无限递归了,和我的数据集配置或者标注文件有关吗?
备注:内容来源于stack exchange,提问作者l3126r
相关产品推荐
相关产品推荐

