You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PaddleOCR自定义训练模型推理出现ValueError问题求助

问题描述

使用自定义数据训练Paddle检测模型后,执行导出命令:

python3 tools/export_model.py -c configs/det/det_r50_vd_db.yml -o Global.pretrained_model="./output/det_r50_vd/latest" Global.save_inference_dir="./output/det_db_inference/"

导出输出内容:

W0804 12:55:34.817917 4102 gpu_resources.cc:61] Please NOTE: device: 0, GPU Compute Capability: 6.0, Driver API Version: 11.0, Runtime API
Version: 10.2
W0804 12:55:34.822103 4102 gpu_resources.cc:91] device: 0, cuDNN Version: 7.6.
[2022/08/04 12:55:35] ppocr INFO: load pretrain successful from ./output/det_r50_vd/best_accuracy
[2022/08/04 12:55:38] ppocr INFO: inference model is saved to ./output/det_db_inference/inference

随后执行推理命令:

python3 tools/infer/predict_det.py --det_algorithm="DB" --det_model_dir="./output/det_db_inference/" --image_dir="../image" --use_gpu=True

出现如下ValueError报错:

Traceback (most recent call last):
File "tools/infer/predict_det.py", line 262, in
text_detector = TextDetector(args)
File "tools/infer/predict_det.py", line 121, in init
args, 'det', logger)
File "/home/user/paddle/PaddleOCR/tools/infer/utility.py", line 317, in create_predictor
predictor = inference.create_predictor(config)
ValueError: (InvalidArgument) The inverse of Fused batch norm variance should be finite. Found nonfinite values! Please check batch_norm_55.w_2
[Hint: Expected std::isfinite(variance_array[i]) == true, but received std::isfinite(variance_array[i]):0 != true:1.] (at /paddle/paddle/fluid/framework/ir/conv_bn_fuse_pass.cc:105)

请问该问题是什么意思,可能由什么原因导致?

问题解析与原因分析

报错含义

该错误说明模型中batch_norm_55.w_2这个Batch Normalization层的方差出现了非有限值(如NaN或无穷大),导致计算方差的倒数时无法得到有效有限数值,进而在模型加载的Conv-BN融合优化阶段触发验证失败。

可能的原因

  • 训练过程异常:
    • 训练数据存在异常样本,比如像素值为NaN/无穷大,或样本分布极端偏离预训练数据,导致BatchNorm层计算均值、方差时出现数值溢出。
    • 训练学习率设置过高,引发模型参数更新幅度过大,导致BatchNorm层的统计值(均值、方差)出现异常。
    • 训练过程中断或异常保存,导致导出的模型参数不完整或损坏。
  • 模型导出问题:
    • 导出命令指定的预训练模型路径与实际加载的不一致:命令中写的是latest,但日志显示加载了best_accuracy,可能存在模型文件不匹配的情况,导致导出的推理模型参数异常。
    • 导出时的Conv-BN融合优化逻辑与训练时的设置不兼容,引发数值计算异常。
  • 硬件/环境差异:
    • 训练和推理时的CUDA、cuDNN版本不匹配(日志显示训练时CUDA Runtime版本为10.2,Driver版本为11.0),可能导致模型参数在不同环境下加载时出现数值异常。

内容的提问来源于stack exchange,提问作者Vikas Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 06:45:59