You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

添加BatchNorm与Dropout后自动编码器损失不收敛问题求助

自编码器训练异常问题求助

研究自编码器已有数周,目前在损失函数理解上遇到瓶颈。在模型中加入BatchNormalization和Dropout层后,出现损失不收敛、重构效果极差的问题,典型损失曲线表现为训练与验证损失始终处于高位且无明显下降趋势。

我采用L1正则化结合MSE损失,相关代码如下:

损失函数代码

def L1_loss_fcn(model_children, true_data, reconstructed_data, reg_param=0.1, validate):
    mse = nn.MSELoss()
    mse_loss = mse(reconstructed_data, true_data)

    l1_loss = 0
    values = true_data
    if validate == False:
        for i in range(len(model_children)):
            values = F.relu((model_children[i](values)))
            l1_loss += torch.sum(torch.abs(values))

        loss = mse_loss + reg_param * l1_loss
        return loss, mse_loss, l1_loss
    else: 
        return mse_loss

训练循环代码

optimizer = torch.optim.Adam(model.parameters(), lr=0.001)
train_run_loss = 0
val_run_loss = 0
for epoch in range(epochs):
    print(f"Epoch {epoch + 1} of {epochs}")
    
    # TRAINING
    model.train()
    for data in tqdm(train_dl):
        x, _ = data
        reconstructions = model(x)
        optimizer.zero_grad()
        train_loss, mse_loss, l1_loss = L1_loss_fcn(model_children=model_children, true_data=x, reg_param=regular_param, 
                                                    reconstructed_data=reconstructions, validate=False)
        train_loss.backward()
        optimizer.step()
        train_run_loss += train_loss.item()
    # VALIDATING 
    model.eval()
    with torch.no_grad():
        for data in tqdm(test_dl):
            x, _ = data
            reconstructions = model(x)
            val_loss = L1_loss_fcn(model_children=model_children, true_data=x, reg_param=regular_param, 
                                  reconstructed_data = reconstructions, validate = True)
            val_run_loss += val_loss.item()

epoch_loss_train = train_run_loss / len(train_dl)
epoch_loss_val = val_run_loss / len(test_dl)                

模型结构代码

encoder = nn.Sequential(nn.Linear(), nn.Dropout(p=0.5), nn.LeakyReLU(), nn.BatchNorm1d(),
                        nn.Linear(), nn.Dropout(p=0.4), nn.LeakyReLU(), nn.BatchNorm1d(),
                        nn.Linear(), nn.Dropout(p=0.3), nn.LeakyReLU(), nn.BatchNorm1d(),
                        nn.Linear(), nn.Dropout(p=0.2), nn.LeakyReLU(), nn.BatchNorm1d(),
)
decoder = nn.Sequential(nn.Linear(), nn.Dropout(p=0.2), nn.LeakyReLU(),
                        nn.Linear(), nn.Dropout(p=0.3), nn.LeakyReLU(), 
                        nn.Linear(), nn.Dropout(p=0.4), nn.LeakyReLU(), 
                        nn.Linear(), nn.Dropout(p=0.5), nn.ReLU(), 
)

已尝试调整多种超参数但均无效,希望解决训练与验证损失收敛问题,提升重构效果,恳请技术帮助!

内容的提问来源于stack exchange,提问作者Boston

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 16:46:28