复现GCN模型时训练与验证准确率无变化问题求助
问题:GCN模型训练/验证准确率始终固定不变
问题背景
复现学术论文中的GCN模型,代码运行无报错,但训练、验证准确率全程固定在0.447左右,尝试多种调整方案均引入新错误,难以排查。
实验数据:基于PyTorch Geometric构建的图数据,已验证有效性。每个图包含18个节点,每个节点有18个特征;图结构稀疏,不同图的边数不固定且边带有权重。
模型代码
class GCNModel(nn.Module): def __init__(self): super(GCNModel, self).__init__() # Block1 self.conv1 = GCNConv(in_channels=18, out_channels=64) # Graph Convolution self.relu1 = nn.ReLU() self.conv2 = GCNConv(in_channels=64, out_channels=32) # Graph Convolution self.relu2 = nn.ReLU() self.pool = global_mean_pool # Block2 self.fc1 = nn.Linear(32, 32) # Fully Connected self.dropout1 = nn.Dropout(0.3) # Dropout layer with p=0.3 self.relu3 = nn.ReLU() self.fc2 = nn.Linear(32, 16) # Fully Connected self.dropout2 = nn.Dropout(0.3) # Dropout layer with p=0.3 self.relu4 = nn.ReLU() self.fc3 = nn.Linear(16, 1) # Fully Connected def forward(self, data): x, edge_index, edge_attr, batch = data.x, data.edge_index, data.edge_attr, data.batch # Block1 x = self.conv1(x, edge_index, edge_attr) # Pass edge_attr to the convolution x = self.relu1(x) # ReLU Graph Convolution x = self.conv2(x, edge_index, edge_attr) # Pass edge_attr to the convolution x = self.relu2(x) # Average pooling along the spatial dimension x = self.pool(x, batch) # Block2 Fully Connected x = self.fc1(x) x = self.dropout1(x) # Apply dropout x = self.relu3(x) # ReLU Fully Connected x = self.fc2(x) x = self.dropout2(x) # Apply dropout x = self.relu4(x) # ReLU Fully Connected x = self.fc3(x) return x
训练与评估代码
# Instantiate GCNModel model = GCNModel() # Define optimizer and learning rate optimizer = optim.Adam(model.parameters(), lr=0.001, weight_decay=0.0001) # Define your loss function criterion = nn.CrossEntropyLoss() # Define the data loaders train_loader = DataLoader(data_train, batch_size=64, shuffle=True) val_loader = DataLoader(data_val, batch_size=64, shuffle=False) test_loader = DataLoader(data_test, batch_size=64, shuffle=False) # Training loop def train(epoch): model.train() for data in train_loader: optimizer.zero_grad() output = model(data) loss = criterion(output, data.y.view(-1, 1).float()) loss.backward() optimizer.step() # Function to compute accuracy def compute_accuracy(loader): model.eval() predictions = [] labels = [] with torch.no_grad(): for data in loader: output = model(data) predictions.extend(torch.argmax(output, axis=1).cpu().numpy()) # Use argmax for predicted labels labels.extend(data.y.cpu().numpy()) accuracy = accuracy_score(labels, predictions) return accuracy # Training and validation num_epochs = 50 # Train for 50 epochs for epoch in range(1, num_epochs + 1): train(epoch) train_acc = compute_accuracy(train_loader) # Compute accuracy on training data val_acc = compute_accuracy(val_loader) print(f"Epoch [{epoch}/{num_epochs}], Train Acc: {train_acc:.4f}, Val Acc: {val_acc:.4f}") # Testing test_acc = compute_accuracy(test_loader) print(f"Testing Accuracy: {test_acc:.4f}")
输出结果
Epoch [1/50], Train Acc: 0.4471, Val Acc: 0.4470 Epoch [2/50], Train Acc: 0.4471, Val Acc: 0.4470 . . . Epoch [50/50], Train Acc: 0.4471, Val Acc: 0.4470
排查思路与解决方案
核心问题定位
准确率固定不变的直接原因是预测结果完全没有变化,本质是模型输出不随训练迭代更新,或预测逻辑错误。
具体修复步骤
修正损失函数与任务匹配度
- 当前用
CrossEntropyLoss但模型最后一层输出为1维,完全不匹配:CrossEntropyLoss要求输出维度为[batch_size, num_classes],且标签为整数类型。 - 二分类场景:改用
BCEWithLogitsLoss,同时调整标签处理:criterion = nn.BCEWithLogitsLoss() # 训练时损失计算 loss = criterion(output.squeeze(), data.y.float()) - 多分类场景:修改模型最后一层输出维度为类别数,比如2类时:
# 模型__init__中修改 self.fc3 = nn.Linear(16, 2) # 损失函数与标签处理 criterion = nn.CrossEntropyLoss() loss = criterion(output, data.y.long())
- 当前用
修正预测逻辑
- 二分类场景:用sigmoid转换后以0.5为阈值判断类别:
predictions.extend((torch.sigmoid(output.squeeze()) > 0.5).int().cpu().numpy()) - 多分类场景:保持argmax,但需确保输出维度为
[batch_size, num_classes]:predictions.extend(torch.argmax(output, axis=1).cpu().numpy())
- 二分类场景:用sigmoid转换后以0.5为阈值判断类别:
验证边权重的有效性
PyTorch Geometric的GCNConv对边权重的处理需要手动归一化,否则边权重可能未生效:from torch_geometric.utils import add_self_loops, degree # 在forward中修改卷积部分 edge_index, edge_attr = add_self_loops(edge_index, edge_attr, fill_value=1.0) row, col = edge_index deg = degree(col, x.size(0), dtype=x.dtype) deg_inv_sqrt = deg.pow(-0.5) deg_inv_sqrt[deg_inv_sqrt == float('inf')] = 0 norm = deg_inv_sqrt[row] * edge_attr * deg_inv_sqrt[col] x = self.conv1(x, edge_index, edge_weight=norm)验证参数更新状态
训练前记录模型参数(比如conv1.weight)的初始值,训练2-3个epoch后对比参数是否变化。如果参数无更新,说明梯度未传递,需检查损失函数计算、反向传播流程是否存在错误。调整优化器配置
暂时移除weight_decay=0.0001,或把学习率调高至0.01,验证模型是否能正常更新参数,排除正则化过强或学习率过低的问题。
内容的提问来源于stack exchange,提问作者Mohammed Nafie
相关产品推荐
相关产品推荐

