You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Style Loss始终为零问题求助:PyTorch风格迁移实现排查

排查风格损失始终为零的问题

我看到你的代码里有几个关键问题导致风格损失一直为零,咱们一步步来分析和修正:

1. 错误的特征提取逻辑

你当前在遍历VGG层时,每次匹配到目标层(比如conv_5或conv_16),都直接调用self.vgg_model(input_)——这会让输入完整跑完整个VGG模型,得到的是最后一层的输出,而不是你指定的目标层的特征。这就导致你计算风格损失时根本没用到conv_5层的特征,而是用了VGG最终层的输出,这大概率会让Gram矩阵的差异消失(或者你的输入/目标经过完整VGG后输出完全一致),最终风格损失为零。

正确做法:逐层传递输入,把输入和目标依次通过每一层,当到达你指定的content/style层时,记录当前的特征。

2. Gram矩阵计算错误

你的Gram矩阵把batch维度和通道维度合并了(features = input_.view(a*b, c*d)),这会计算跨样本的通道内积,而不是每个样本内部通道之间的内积——这完全违背了风格损失中Gram矩阵的定义,也会导致损失计算异常。

正确做法:针对每个样本,将特征展平为(batch_size, channels, height*width),然后计算通道间的内积,得到(batch_size, channels, channels)的Gram矩阵。

3. 冗余代码问题

你的gram_matrix函数里有两个return语句,第二个return G永远不会被执行,虽然不影响功能,但需要清理。


修正后的完整代码

import torch
import torch.nn as nn
import torch.nn.functional as F
from torchvision import models

class Feature_and_style_losses():
    def __init__(self):
        self.vgg_model = models.vgg19(pretrained=True).features.cuda().eval()
        # 固定VGG参数,避免反向传播更新
        for param in self.vgg_model.parameters():
            param.requires_grad = False
        self.content_layers = ['conv_16']
        self.style_layers = ['conv_5']

    def calculate_feature_and_style_losses(self, input_, target, feature_coefficient, style_coefficient):
        i = 0
        feature_losses = []
        style_losses = []
        # 初始化当前特征,逐层传递
        input_features = input_.clone()
        target_features = target.clone()

        for layer_ in self.vgg_model.children():
            # 逐层处理输入和目标特征
            input_features = layer_(input_features)
            target_features = layer_(target_features)

            if isinstance(layer_, nn.Conv2d):
                i += 1
                name = f"conv_{i}"
                # 计算内容损失
                if name in self.content_layers:
                    feat_loss = self.feature_loss(input_features, target_features)
                    feature_losses.append(feat_loss)
                # 计算风格损失
                if name in self.style_layers:
                    style_loss = self.style_loss(input_features, target_features)
                    style_losses.append(style_loss)
            # 如果是池化层或ReLU层,不计数conv,直接继续
            elif isinstance(layer_, nn.MaxPool2d) or isinstance(layer_, nn.ReLU):
                continue

        # 计算总损失
        feature_loss_value = torch.mean(torch.tensor(feature_losses)) * feature_coefficient
        style_loss_value = torch.mean(torch.tensor(style_losses)) * style_coefficient
        return feature_loss_value, style_loss_value

    def feature_loss(self, input_, target):
        target = target.detach()
        return F.mse_loss(input_, target)

    def gram_matrix(self, input_):
        batch_size, channels, height, width = input_.size()
        # 展平为 (batch_size, channels, height*width)
        features = input_.view(batch_size, channels, height * width)
        # 计算Gram矩阵:(batch_size, channels, channels)
        G = torch.bmm(features, features.transpose(1, 2))
        # 归一化,除以特征元素总数
        return G.div(channels * height * width)

    def style_loss(self, input_, target):
        G_input = self.gram_matrix(input_)
        G_target = self.gram_matrix(target).detach()
        return F.mse_loss(G_input, G_target)

额外注意事项

  • 我添加了for param in self.vgg_model.parameters(): param.requires_grad = False,固定VGG的预训练参数,避免损失反向传播时更新VGG权重,这是神经风格迁移中的标准做法。
  • 修正了遍历层时的处理逻辑,确保只有Conv2d层才会计数命名,池化和ReLU层跳过计数。
  • 用torch.tensor替代了torch.from_numpy(np.array(...)),避免不必要的CPU-GPU数据传输,更高效。

内容的提问来源于stack exchange,提问作者Can

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 21:12:49