使用相同权重与网络为何推理结果存在差异?
背景
我需要把YOLOv5的预训练层权重用到自定义多任务网络中,因此先保存了原YOLOv5模型的state_dict,再在自定义模型里手动覆盖目标检测相关层的权重(因层名不一致,无法直接调用model.load_state_dict),推理使用自定义detect模块。
问题现象
使用完全相同的权重与输入,两个模型的推理结果差异极大。对比中间层输出发现,数据在网络传递过程中差异逐渐扩大,虽模型与输入均为32位浮点,但疑似微小浮点误差经多层累积后导致最终结果偏差。
原模型权重保存代码
在YOLOv5项目的detect.py推理前执行:
torch.save(model.model.state_dict(), 'yolov5m_state_dict.pt')
自定义模型定义(省略其他任务头)
import torch.nn as nn import torch # 假设Conv、C3、SPPF、Detect为YOLOv5原模块的复用 class YOLOPointMFull(nn.Module): def __init__(self, inp_ch=3, nc=80, anchors=None): super(YOLOPointMFull, self).__init__() if anchors is None: anchors = [] # 沿用YOLOv5m的anchors配置 # CSPNet Backbone self.Conv1 = Conv(inp_ch, 48, 6, 2, 2) # ch_in, ch_out, kernel, stride, padding, groups self.Conv2 = Conv(48, 96, 3, 2) self.Bottleneck1 = C3(96, 96, 2) # ch_in, ch_out, number self.Conv3 = Conv(96, 192, 3, 2) self.Bottleneck2 = C3(192, 192, 4) self.Conv4 = Conv(192, 384, 3, 2) self.Bottleneck3 = C3(384, 384, 6) self.Conv5 = Conv(384, 768, 3, 2) self.Bottleneck4 = C3(768, 768, 2) self.SPPooling = SPPF(768, 768, 5) # Object Detector Head self.Conv6 = Conv(768, 384, 1, 1, 0) # ups, cat self.Bottleneck5 = C3(768, 384, 2) self.Conv7 = Conv(384, 192, 1, 1, 0) # ups, cat self.Bottleneck6 = C3(384, 192, 2) # --> detect self.Conv8 = Conv(192, 192, 3, 2, 1) # cat self.Bottleneck7 = C3(384, 384, 2) # --> detect self.Conv9 = Conv(384, 384, 3, 2, 1) # cat self.Bottleneck8 = C3(768, 768, 2) # --> detect self.Detect = Detect(nc, anchors=anchors, ch=(192, 384, 768)) self.ups = nn.Upsample(scale_factor=(2, 2), mode='nearest') def forward(self, x): # Backbone x = self.Conv1(x) x = self.Conv2(x) # check x = self.Bottleneck1(x) x = self.Conv3(x) xb = self.Bottleneck2(x) x = self.Conv4(xb) xc = self.Bottleneck3(x) x = self.Conv5(xc) x = self.Bottleneck4(x) x = self.SPPooling(x) # Object Detector Head xd = self.Conv6(x) x = self.ups(xd) x = torch.cat((x, xc), dim=1) x = self.Bottleneck5(x) xe = self.Conv7(x) x = self.ups(xe) x = torch.cat((x, xb), dim=1) xf = self.Bottleneck6(x) x = self.Conv8(xf) x = torch.cat((x, xe), dim=1) xg = self.Bottleneck7(x) x = self.Conv9(xg) x = torch.cat((x, xd), dim=1) x = self.Bottleneck8(x) x = self.Detect([xf, xg, x]) return x
手动覆盖权重的代码
yolo_sd = torch.load('yolov5m_state_dict.pt') # 实例化自定义模型,使用默认anchors、输入通道和类别数 model = YOLOPointMFull() yp_full_sd = model.state_dict() def load_pretrained(target, source): print(len(target), len(source)) for i, (tk, sk) in enumerate(zip(target, source)): layer_target = tk.split('.')[-1] layer_source = sk.split('.')[-1] if layer_source == layer_target and source[sk].shape == target[tk].shape: target[tk] = source[sk] else: print(f'Failed to overwrite {tk}') return target yp_full_sd = load_pretrained(yp_full_sd, yolo_sd) torch.save({"model_state_dict": yp_full_sd,}, 'logs/full_model/yolo_pure.pt')
已排查项
- 输入完全一致:RGB格式、尺寸、float32类型、像素值无差异,无额外预处理
- 尝试过不同权重保存格式(.pt、.pth.tar),结果一致
- 原
state_dict来自未融合的YOLO模型(BN层与Conv层分离) - 推理时模型均已设置为
eval模式 - 两个模型在同一GPU运行
- 非极大值抑制(NMS)参数完全相同
- 环境差异:
- 自定义项目:torch 1.11.0/1.12.1、Python 3.8.13、numpy 1.23.1、opencv-python 4.5.5.64
- YOLOv5项目:torch 1.10.0+cu113、Python 3.8.0、opencv-python 4.5.4.58
补充定位
用torch.ones(1, 3, 256, 256, dtype=torch.float)作为输入测试,发现差异始于第一个C3模块的ConvBNSiLU层之后,首次差异的张量平均绝对差仅为9.7185235e-05,属于微小浮点误差,且已确认Conv层权重完全一致。
解决思路
统一环境版本:两个项目PyTorch版本不一致(1.10 vs 1.11/1.12),不同版本对算子(如BN层、SiLU激活)的浮点计算逻辑可能存在细微差异。建议将自定义项目的PyTorch版本降级到与YOLOv5一致的1.10.0+cu113,排除版本差异导致的计算偏差。
修正权重加载逻辑:当前按遍历顺序匹配层名后缀的方式容易出现权重赋值错误。原YOLOv5模型层按
model.model.0、model.model.1索引命名,自定义模型按Conv1、Bottleneck1命名,需手动建立模块功能映射关系,示例:layer_mapping = { "model.0": "Conv1", "model.1": "Conv2", "model.2": "Bottleneck1", # 补全所有对应层的映射 } for src_key_prefix, target_key_prefix in layer_mapping.items(): # 处理权重、偏置 for param_suffix in [".weight", ".bias"]: src_key = src_key_prefix + param_suffix target_key = target_key_prefix + param_suffix if src_key in yolo_sd and target_key in yp_full_sd: if yolo_sd[src_key].shape == yp_full_sd[target_key].shape: yp_full_sd[target_key] = yolo_sd[src_key] else: print(f"Shape mismatch: {src_key} -> {target_key}") # 处理BN层的running_mean、running_var for bn_suffix in [".bn.running_mean", ".bn.running_var", ".bn.weight", ".bn.bias"]: src_key = src_key_prefix + bn_suffix target_key = target_key_prefix + bn_suffix if src_key in yolo_sd and target_key in yp_full_sd: yp_full_sd[target_key] = yolo_sd[src_key]检查模块实现一致性:确保自定义模型中复用的
Conv、C3、SPPF、Detect模块与YOLOv5原代码完全一致,重点确认:Conv模块的执行顺序:Conv -> BN -> SiLU- C3模块的内部残差连接逻辑
- SiLU激活的实现(直接复制YOLOv5原代码中的SiLU实现到自定义项目,避免版本差异)
强制确定性计算:推理前设置以下参数,避免CuDNN非确定性算法导致的浮点差异:
torch.backends.cudnn.deterministic = True torch.backends.cudnn.benchmark = False校验BN层参数:确认原模型BN层的
running_mean、running_var、weight、bias均被正确加载,当前匹配逻辑可能遗漏这些参数,导致BN层计算不一致。
内容的提问来源于stack exchange,提问作者Toni

