使用PyTorch导出.pt模型到ONNX时遇AttributeError问题求助
解决.pt转ONNX时的AttributeError: 'OrderedDict' object has no attribute 'modules'错误
错误原因
你通过torch.load()加载的base_40m_textvec.pt文件,保存的是模型的状态字典(OrderedDict),而非完整的PyTorch模型实例。torch.onnx.export()要求传入可调用的模型对象,而非仅包含权重的状态字典,因此触发了属性缺失错误。
解决方案
1. 还原模型结构
状态字典仅存储模型权重,你需要先定义和训练时完全一致的模型类结构。示例如下(请替换为你实际的模型结构):
import torch import torch.nn as nn class YourTextVecModel(nn.Module): def __init__(self): super().__init__() # 此处编写你模型的实际层定义,比如卷积、Transformer层等 self.conv1 = nn.Conv2d(3, 64, kernel_size=3, padding=1) self.fc = nn.Linear(64*224*224, 128) # ...其他层定义 def forward(self, x): # 编写模型的前向传播逻辑 x = self.conv1(x) x = x.flatten(start_dim=1) x = self.fc(x) return x
2. 加载状态字典到模型实例
实例化模型后,将状态字典加载到模型中:
# 实例化模型并移至CUDA设备 model = YourTextVecModel().to("cuda") # 加载状态字典 state_dict = torch.load('base_40m_textvec.pt') # 将权重加载到模型 model.load_state_dict(state_dict) # 切换到评估模式(推荐,避免BatchNorm、Dropout等层影响导出结果) model.eval()
3. 重新执行ONNX导出
完成上述步骤后,再执行你的导出代码即可:
dummy_input = torch.randn(10, 3, 224, 224, device="cuda") input_names = [ "actual_input_1" ] + [ "learned_%d" % i for i in range(16) ] output_names = [ "output1" ] torch.onnx.export(model, dummy_input, "base_40m_textvec.onnx", verbose=True, input_names=input_names, output_names=output_names)
关键注意事项
- 模型结构必须和训练时完全一致,包括层的数量、参数维度、命名规则,否则
load_state_dict会抛出不匹配的错误。 - 如果训练时是用
torch.save(model, path)保存的完整模型,直接model = torch.load(path)即可,但你的文件仅保存了状态字典,因此必须按上述流程操作。
内容的提问来源于stack exchange,提问作者tempaccountforissue
相关产品推荐
相关产品推荐

