You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VGG19风格迁移PyTorch模型转ONNX输出维度异常问题

问题背景
  • 参考开源神经风格迁移项目实现相关功能,训练完成后计划将模型导出为ONNX格式,用于C++搭配OpenCV环境部署。
  • 导出的ONNX模型存在输出维度异常:输入维度为(1,3,512,512)时,输出维度为(1,512,28,28),不符合部署要求。

使用的导出代码如下:

VGG.eval()
torch.save(VGG, 'torchmodel.pth')
dummy_input = Variable(torch.randn(1, 3, 512, 512, device='cuda:1'))
input_names = ['input']
output_names = ['output']
onnxfile='style.onnx'
torch.onnx.export(VGG,dummy_input,onnxfile,verbose=False,input_names=input_names,opset_version=11, output_names=output_names)

项目原始实现代码如下:

# -*- coding: utf-8 -*-
"""StyleTransfer.ipynb
Automatically generated by Colaboratory.
"""

import torch
import torch.nn as nn
import torch.optim as optim
import torchvision as tv
from PIL import Image
import imageio
import numpy as np
from matplotlib import pyplot as plt

to_tensor = tv.transforms.Compose([
            tv.transforms.Resize((512,512)),
            tv.transforms.ToTensor(),
            tv.transforms.Normalize(mean=[0.485, 0.456, 0.406],
                                std=[1, 1, 1]),
        ])

unload = tv.transforms.Compose([
            tv.transforms.Normalize(mean=[-0.485,-0.456,-0.406],
                                std=[1,1,1]),                
            tv.transforms.Lambda(lambda x: x.clamp(0,1))
        ])
to_image = tv.transforms.ToPILImage()

style_img = 'udnie.jpg'
input_img = 'chicago.jpg'

style_img = Image.open(style_img)
input_img = Image.open(input_img)

style_img = to_tensor(style_img).cuda()
input_img = to_tensor(input_img).cuda()

def get_features(module, x, y):
    features.append(y)

def gram_matrix(x):
    b, c, h, w = x.size()
    F = x.view(b,c,h*w)
    G = torch.bmm(F, F.transpose(1,2))/(h*w)
    return G

VGG = tv.models.vgg19(pretrained=True).features
VGG.cuda()

for i, layer in enumerate(VGG):
    if i in [0,5,10,19,21,28]:
        VGG[i].register_forward_hook(get_features)
    elif isinstance(layer, nn.MaxPool2d):
        VGG[i] = nn.AvgPool2d(kernel_size=2)

VGG.eval()

for p in VGG.parameters():
    p.requires_grad = False

features = []
VGG(input_img.unsqueeze(0))
c_target = features[4].detach()

features = []
VGG(style_img.unsqueeze(0))
f_targets = features[:4]+features[5:]
gram_targets = [gram_matrix(i).detach() for i in f_targets]

alpha = 1
beta = 1e3
iterations = 200
image = input_img.clone().unsqueeze(0)
images = []
optimizer = optim.LBFGS([image.requires_grad_()], lr=1)    
mse_loss = nn.MSELoss(reduction='mean')
l_c = []
l_s = []
counter = 0

for itr in range(iterations):
    features = []
    def closure():
        optimizer.zero_grad()
        VGG(image)
        t_features = features[-6:]
        content = t_features[4]
        style_features = t_features[:4]+t_features[5:]
        t_features = []
        gram_styles = [gram_matrix(i) for i in style_features]
        c_loss = alpha * mse_loss(content, c_target)
        s_loss = 0
        for i in range(5):
            n_c = gram_styles[i].shape[0]
            s_loss += beta * mse_loss(gram_styles[i],gram_targets[i])/(n_c**2)
        total_loss = c_loss+s_loss
        l_c.append(c_loss)
        l_s.append(s_loss)
        total_loss.backward()
        return total_loss

    optimizer.step(closure)
    print('Step {}: S_loss: {:.8f} C_loss: {:.8f}'.format(itr, l_s[-1], l_c[-1]))

    if itr%1 == 0:
        temp = unload(image[0].cpu().detach())
        temp = to_image(temp)
        temp = np.array(temp)
        images.append(temp)
        imageio.mimsave('progress.gif', images)
    
plt.clf()
plt.plot(l_c, label='Content Loss')
plt.legend()
plt.savefig('loss1.png')

plt.clf()
plt.plot(l_s, label='Style Loss')
plt.legend()
plt.savefig('loss2.png')

plt.imsave('last.jpg',images[-1])
问题成因

核心原因是 你导出的模型根本不是用于推理的风格迁移生成模型:

  1. 你当前运行的是Gatys提出的经典基于优化的神经风格迁移算法,逻辑是固定预训练VGG的参数,把待生成的风格化图像作为可优化变量,通过VGG提取中间层特征计算内容损失、风格损失,迭代更新输入图像像素得到最终结果,整个过程不存在“训练好的端到端生成网络”。
  2. 你代码中保存、导出的VGG是仅用于提取特征计算损失的backbone网络,本身只负责输出高层卷积特征,不会输出和输入同尺寸的3通道RGB图像。你得到的(1,512,28,28)输出就是VGG深层卷积的中间特征图,属于特征提取网络的正常输出,和风格迁移的最终生成结果没有关系。
  3. 导出代码中使用的Variable是PyTorch早版本已弃用的写法,不会影响模型输出维度,不是问题诱因。
修复方案
  • 更换为前向式快速风格迁移方案:如果需要单模型前向推理部署,必须使用带独立图像转换网络的快速风格迁移架构,架构分为两部分:一部分是固定权重的VGG损失网络(仅训练时用,部署不需要),另一部分是由下采样层、残差块、上采样层组成的图像转换网络(输入3通道原图,直接输出3通道风格化结果,和输入尺寸一致)。训练完成后仅需要导出图像转换网络部分为ONNX,即可满足C+++OpenCV的部署要求。
  • 不建议硬适配当前算法:如果要基于现有优化式算法部署,需要把VGG前向、Gram矩阵计算、LBFGS优化迭代、损失反传、图像后处理全流程用C++重写,不存在单次前向出结果的可能,推理速度极慢、部署成本极高,无实际落地价值。
  • 导出ONNX前增加校验步骤:每次导出前先在PyTorch环境中执行一次前向,打印输出张量的shape,确认输出维度为预期的(1,3, H, W)后再执行导出操作,避免将损失网络、中间特征提取模块误作为最终推理模型导出。

内容的提问来源于stack exchange,提问作者wiseman s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 13:24:22