You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

单独运行CPU/GPU推理结果正常,同时运行结果异常的原因排查

PyTorch中GPU/CPU混合运行时卷积输出均值不一致问题分析

复现代码

import torch.nn as nn
import torch

torch.manual_seed(200)
torch.cuda.manual_seed(200)
torch.cuda.manual_seed_all(200)

input_v = torch.randn(160, 12, 320, 320)

def result_cuda():
    input = input_v.cuda()
    print("gpu shape = ", input.shape, " mean = ", input.mean().cpu().detach())
    model = nn.Conv2d(12, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False)
    model = model.cuda()
    out = model(input)
    print("gpu shape = ", out.shape, " mean = ", out.mean().cpu().detach())

def result_cpu():
    input = input_v
    print("cpu shape = ", input.shape, " mean = ", input.mean().detach())
    model = nn.Conv2d(12, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False)
    out = model(input)
    print("cpu shape = ", out.shape, " mean = ", out.mean().detach())

result_cuda()
result_cpu()

不同运行场景的输出

单独运行result_cuda()的输出

gpu shape = torch.Size([160, 12, 320, 320]) mean =
tensor(2.1668e-05)

gpu shape = torch.Size([160, 32, 320, 320]) mean =
tensor(7.5935e-06)

单独运行result_cpu()的输出

cpu shape = torch.Size([160, 12, 320, 320]) mean =
tensor(2.1668e-05)

cpu shape = torch.Size([160, 32, 320, 320]) mean =
tensor(7.5936e-06)

同时运行result_cuda()和result_cpu()的输出

gpu shape = torch.Size([160, 12, 320, 320]) mean =
tensor(2.1668e-05)

gpu shape = torch.Size([160, 32, 320, 320]) mean =
tensor(7.5935e-06)

cpu shape = torch.Size([160, 12, 320, 320]) mean =
tensor(2.1668e-05)

cpu shape = torch.Size([160, 32, 320, 320]) mean =
tensor(-1.1677e-05)

问题原因

核心是CPU随机数生成器的状态被提前消耗,导致后续CPU上初始化的卷积层权重完全不同:

  • PyTorch中nn.Conv2d的权重默认在CPU上初始化,无论之后是否将模型移到GPU。
  • 单独运行result_cuda()时,函数内创建卷积层会消耗CPU随机数生成器的初始状态;当你接着运行result_cpu()时,CPU随机数生成器已经不是初始的种子状态,新创建的卷积层会用剩余的随机数初始化权重,自然和单独运行时的权重不一样,最终输出均值差异巨大。
  • 单独运行result_cpu()时,CPU随机数生成器从初始种子开始,初始化的卷积权重和result_cuda()单独运行时的权重一致,所以输出均值仅存在微小浮点数精度差异。

解决方法

如果要保证不同场景下卷积层初始化的权重一致,只需在每个函数开头重置CPU随机种子:

def result_cuda():
    torch.manual_seed(200)  # 重置CPU随机种子
    input = input_v.cuda()
    print("gpu shape = ", input.shape, " mean = ", input.mean().cpu().detach())
    model = nn.Conv2d(12, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False)
    model = model.cuda()
    out = model(input)
    print("gpu shape = ", out.shape, " mean = ", out.mean().cpu().detach())

def result_cpu():
    torch.manual_seed(200)  # 重置CPU随机种子
    input = input_v
    print("cpu shape = ", input.shape, " mean = ", input.mean().detach())
    model = nn.Conv2d(12, 32, kernel_size=(3, 3), stride=(1, 1), padding=(1, 1), bias=False)
    out = model(input)
    print("cpu shape = ", out.shape, " mean = ", out.mean().detach())

内容的提问来源于stack exchange,提问作者xiaoming

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 03:00:10