You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

复现GCN模型时训练与验证准确率无变化问题求助

问题:GCN模型训练/验证准确率始终固定不变

问题背景

复现学术论文中的GCN模型,代码运行无报错,但训练、验证准确率全程固定在0.447左右,尝试多种调整方案均引入新错误,难以排查。

实验数据:基于PyTorch Geometric构建的图数据,已验证有效性。每个图包含18个节点,每个节点有18个特征;图结构稀疏,不同图的边数不固定且边带有权重。

模型代码

class GCNModel(nn.Module):
    def __init__(self):
        super(GCNModel, self).__init__()
        # Block1
        self.conv1 = GCNConv(in_channels=18, out_channels=64)  # Graph Convolution
        self.relu1 = nn.ReLU()
        self.conv2 = GCNConv(in_channels=64, out_channels=32)  # Graph Convolution
        self.relu2 = nn.ReLU()
        self.pool = global_mean_pool

        # Block2
        self.fc1 = nn.Linear(32, 32)     # Fully Connected
        self.dropout1 = nn.Dropout(0.3)  # Dropout layer with p=0.3
        self.relu3 = nn.ReLU()

        self.fc2 = nn.Linear(32, 16)     #  Fully Connected
        self.dropout2 = nn.Dropout(0.3)  # Dropout layer with p=0.3
        self.relu4 = nn.ReLU()

        self.fc3 = nn.Linear(16, 1)       # Fully Connected

    def forward(self, data):
        x, edge_index, edge_attr, batch = data.x, data.edge_index, data.edge_attr, data.batch

        # Block1
        x = self.conv1(x, edge_index, edge_attr)  # Pass edge_attr to the convolution
        x = self.relu1(x)

        # ReLU Graph Convolution
        x = self.conv2(x, edge_index, edge_attr)  # Pass edge_attr to the convolution
        x = self.relu2(x)

        # Average pooling along the spatial dimension
        x = self.pool(x, batch)

        # Block2 Fully Connected
        x = self.fc1(x)
        x = self.dropout1(x)  # Apply dropout
        x = self.relu3(x)

        # ReLU Fully Connected
        x = self.fc2(x)
        x = self.dropout2(x)  # Apply dropout
        x = self.relu4(x)

        # ReLU Fully Connected
        x = self.fc3(x)

        return x

训练与评估代码

# Instantiate GCNModel
model = GCNModel()

# Define optimizer and learning rate
optimizer = optim.Adam(model.parameters(), lr=0.001, weight_decay=0.0001)

# Define your loss function
criterion = nn.CrossEntropyLoss()

# Define the data loaders
train_loader = DataLoader(data_train, batch_size=64, shuffle=True)
val_loader = DataLoader(data_val, batch_size=64, shuffle=False)
test_loader = DataLoader(data_test, batch_size=64, shuffle=False)

# Training loop
def train(epoch):
    model.train()
    for data in train_loader:
        optimizer.zero_grad()
        output = model(data)
        loss = criterion(output, data.y.view(-1, 1).float())
        loss.backward()
        optimizer.step()

# Function to compute accuracy
def compute_accuracy(loader):
    model.eval()
    predictions = []
    labels = []
    with torch.no_grad():
        for data in loader:
            output = model(data)
            predictions.extend(torch.argmax(output, axis=1).cpu().numpy())  # Use argmax for predicted labels
            labels.extend(data.y.cpu().numpy())
    accuracy = accuracy_score(labels, predictions)
    return accuracy

# Training and validation
num_epochs = 50  # Train for 50 epochs
for epoch in range(1, num_epochs + 1):
    train(epoch)
    train_acc = compute_accuracy(train_loader)  # Compute accuracy on training data
    val_acc = compute_accuracy(val_loader)
    print(f"Epoch [{epoch}/{num_epochs}], Train Acc: {train_acc:.4f}, Val Acc: {val_acc:.4f}")

# Testing
test_acc = compute_accuracy(test_loader)
print(f"Testing Accuracy: {test_acc:.4f}")

输出结果

Epoch [1/50], Train Acc: 0.4471, Val Acc: 0.4470
Epoch [2/50], Train Acc: 0.4471, Val Acc: 0.4470
.
.
.
Epoch [50/50], Train Acc: 0.4471, Val Acc: 0.4470

排查思路与解决方案

核心问题定位

准确率固定不变的直接原因是预测结果完全没有变化,本质是模型输出不随训练迭代更新,或预测逻辑错误。

具体修复步骤

  1. 修正损失函数与任务匹配度

    • 当前用CrossEntropyLoss但模型最后一层输出为1维,完全不匹配:CrossEntropyLoss要求输出维度为[batch_size, num_classes],且标签为整数类型。
    • 二分类场景:改用BCEWithLogitsLoss,同时调整标签处理:
      criterion = nn.BCEWithLogitsLoss()
      # 训练时损失计算
      loss = criterion(output.squeeze(), data.y.float())
      
    • 多分类场景:修改模型最后一层输出维度为类别数,比如2类时:
      # 模型__init__中修改
      self.fc3 = nn.Linear(16, 2)
      # 损失函数与标签处理
      criterion = nn.CrossEntropyLoss()
      loss = criterion(output, data.y.long())
      
  2. 修正预测逻辑

    • 二分类场景:用sigmoid转换后以0.5为阈值判断类别:
      predictions.extend((torch.sigmoid(output.squeeze()) > 0.5).int().cpu().numpy())
      
    • 多分类场景:保持argmax,但需确保输出维度为[batch_size, num_classes]:
      predictions.extend(torch.argmax(output, axis=1).cpu().numpy())
      
  3. 验证边权重的有效性
    PyTorch Geometric的GCNConv对边权重的处理需要手动归一化,否则边权重可能未生效:

    from torch_geometric.utils import add_self_loops, degree
    
    # 在forward中修改卷积部分
    edge_index, edge_attr = add_self_loops(edge_index, edge_attr, fill_value=1.0)
    row, col = edge_index
    deg = degree(col, x.size(0), dtype=x.dtype)
    deg_inv_sqrt = deg.pow(-0.5)
    deg_inv_sqrt[deg_inv_sqrt == float('inf')] = 0
    norm = deg_inv_sqrt[row] * edge_attr * deg_inv_sqrt[col]
    x = self.conv1(x, edge_index, edge_weight=norm)
    
  4. 验证参数更新状态
    训练前记录模型参数(比如conv1.weight)的初始值,训练2-3个epoch后对比参数是否变化。如果参数无更新,说明梯度未传递,需检查损失函数计算、反向传播流程是否存在错误。

  5. 调整优化器配置
    暂时移除weight_decay=0.0001,或把学习率调高至0.01,验证模型是否能正常更新参数,排除正则化过强或学习率过低的问题。

内容的提问来源于stack exchange,提问作者Mohammed Nafie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:32:16