You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

灵活神经网络模型Loss曲线平坦问题排查求助

问题描述

搭建的灵活神经网络模型出现Loss曲线平坦(无下降趋势)的问题,此前使用无动态层数配置的常规神经网络模型时,Loss能够正常下降,请求排查原因。

相关代码与配置

模型代码

class FlexibleModel(nn.Module):
    def __init__(self, in_features, hidden_layers_sizes, output_size=1):
        super().__init__()
        self.layers = nn.ModuleList() # Use ModuleList to store layers

        current_in_features = in_features
        for h_size in hidden_layers_sizes:
            self.layers.append(nn.Linear(current_in_features, h_size))
            current_in_features = h_size # Output of current layer becomes input for next

        self.out = nn.Linear(current_in_features, output_size)

    def forward(self, x):
        for layer in self.layers:
            x = F.relu(layer(x)) # Apply ReLU after each hidden layer
        x = self.out(x)
        return x

学习率与配置

hidden_config_1 = [128, 64]
model_1 = FlexibleModel(in_features=780, hidden_layers_sizes=hidden_config_1)
print("Model 1 Architecture:")
print(model_1)

optimizer = torch.optim.Adam(model_1.parameters(), lr=0.001)
criterion = nn.MSELoss()

训练循环

num_epochs = 1000
loss_list1 = []

for epoch in range(num_epochs):
    model_1.train()

    # Forward pass
    y_pred = model_1(X_train)
    loss = F.mse_loss(y_pred, y_train)

    # Backward + optimize
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

    loss_list1.append(loss.item())

    # Validation
    model_1.eval()
    with torch.no_grad():
        val_pred = model_1(X_val)
        val_loss = F.mse_loss(val_pred, y_val)
排查方向建议
  • 参数初始化优化:ReLU激活配合Kaiming初始化能缓解梯度消失,默认线性层的均匀初始化可能导致部分神经元输出为0且无法恢复。可修改模型初始化代码,给每个Linear层添加Kaiming初始化:
    # 在__init__的循环里替换原添加层的代码
    layer = nn.Linear(current_in_features, h_size)
    nn.init.kaiming_normal_(layer.weight, mode='fan_in', nonlinearity='relu')
    nn.init.zeros_(layer.bias)
    self.layers.append(layer)
    
  • 输入输出维度校验:检查X_train、y_train的维度是否匹配模型输出。例如模型输出是(batch_size, 1),若y_train是一维张量(batch_size,),可能触发广播机制导致Loss计算异常,需将y_train转为二维:y_train = y_train.unsqueeze(1)。
  • 数据归一化:输入特征维度达780,若各特征数值范围差异过大,会导致模型难以收敛。对输入数据做标准化(如StandardScaler)或归一化,确保特征分布在相近区间。
  • 梯度值检查:在训练循环中添加梯度监控,查看是否存在梯度消失:
    loss.backward()
    # 打印第一层权重的梯度范数
    print(model_1.layers[0].weight.grad.norm())
    
    若梯度范数接近0,说明梯度消失;若过大则是梯度爆炸,需调整初始化或添加梯度裁剪。
  • Loss计算一致性:统一使用criterion(y_pred, y_train)代替F.mse_loss,避免因参数顺序或 dtype 不一致导致的问题,同时确保y_pred和y_train的 dtype 相同(如均为float32)。
  • 优化器参数校验:确认optimizer.param_groups中的参数数量与模型可训练参数数量一致,避免出现优化器未绑定全部参数的情况。

内容的提问来源于stack exchange,提问作者tommy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 22:00:14