基于GGN的SWIFT网络欺诈检测:边分类维度越界问题求助
问题:GCN边分类时出现索引越界错误
问题背景
我正在用GCN进行SWIFT网络中的欺诈交易检测,图中以银行SWIFT BIC码为节点,边代表交易。边标签0表示欺诈交易、1表示正常交易,每条边包含金额、事由、货币、时间戳4个特征。各张量形状如下:
- 节点特征张量:
torch.Size([1, 210, 6]) - 边特征张量:
torch.Size([200, 4]) - 边索引张量:
torch.Size([2, 200]) - 邻接矩阵张量:
torch.Size([1, 210, 210]) - 标签张量:
torch.Size([200, 1])
原实现代码
import torch import torch.nn as nn import torch.nn.functional as F from torch_geometric.nn import GCNConv class GCNEdgeClassifier(nn.Module): def __init__(self, input_dim, hidden_dim, output_dim): super(GCNEdgeClassifier, self).__init__() self.conv1 = GCNConv(input_dim, hidden_dim) self.conv2 = GCNConv(hidden_dim, output_dim) def forward(self, x, edge_index): # x: edge features (attributes) # edge_index: edge indexes # Apply first graph convolutional layer print("X ", x.shape) print("edge_index ",edge_index.shape) x = self.conv1(x, edge_index) x = F.relu(x) x = F.dropout(x, p=0.5, training=self.training) # Apply second graph convolutional layer x = self.conv2(x, edge_index) # Perform edge classification # Here we can use a linear layer, sigmoid, or any other activation x = torch.sigmoid(x) return x # Define the model input_dim = 4 # Number of edge features hidden_dim = 64 # Hidden dimension size output_dim = 1 # Output dimension for edge classification (binary) model = GCNEdgeClassifier(input_dim, hidden_dim, output_dim) # Forward pass predictions = model(edge_features_tensor, edge_indexes_tensor) # Loss calculation (assuming binary cross-entropy loss) loss_fn = nn.BCELoss() labels_tensor = labels_tensor.float() # Convert labels to float for BCELoss loss = loss_fn(predictions, labels_tensor) # Backpropagation and parameter optimization (if training) loss.backward() # optimizer.step() # Update parameters if using an optimizer
报错信息
RuntimeError: index 202 is out of bounds for dimension 0 with size 200
说明:共有200条边(对应200个标签和边特征),但边索引引用了210个节点,刚接触GCN,不清楚如何解决边分类的适配问题。
错误原因
你使用的GCNConv(PyTorch Geometric中的标准层)默认是基于节点特征做图卷积,它要求输入的x是节点特征张量(形状为[节点数, 特征数]),但你传入的是边特征(仅200个样本)。而边索引中的节点索引(比如202)指向的是第203个节点(索引从0开始),此时输入的x只有200个元素,自然会触发索引越界。
修正方案
边分类需要结合节点的上下文信息与边自身特征,以下是可行的修正代码:
import torch import torch.nn as nn import torch.nn.functional as F from torch_geometric.nn import GCNConv class GCNEdgeClassifier(nn.Module): def __init__(self, node_input_dim, hidden_dim, output_dim): super(GCNEdgeClassifier, self).__init__() # GCNConv输入维度为节点特征维度(6) self.conv1 = GCNConv(node_input_dim, hidden_dim) self.conv2 = GCNConv(hidden_dim, hidden_dim) # 通过边两端节点特征拼接做分类,可选择加入边特征 self.fc = nn.Linear(hidden_dim * 2 + 4, output_dim) def forward(self, node_x, edge_index, edge_x): # node_x: 节点特征,形状[210, 6](需去掉batch维度) # edge_index: 边索引,形状[2, 200] # edge_x: 边特征,形状[200, 4] # 节点卷积层提取节点上下文特征 node_x = self.conv1(node_x, edge_index) node_x = F.relu(node_x) node_x = F.dropout(node_x, p=0.5, training=self.training) node_x = self.conv2(node_x, edge_index) # 获取每条边的源节点与目标节点特征 src_node_feat = node_x[edge_index[0]] # 形状[200, hidden_dim] dst_node_feat = node_x[edge_index[1]] # 形状[200, hidden_dim] # 拼接节点特征与边特征,生成边的表示 edge_repr = torch.cat([src_node_feat, dst_node_feat, edge_x], dim=1) # 边分类预测 out = self.fc(edge_repr) out = torch.sigmoid(out) return out # 初始化模型:节点特征维度6,隐藏层64,输出维度1 model = GCNEdgeClassifier(node_input_dim=6, hidden_dim=64, output_dim=1) # 处理节点特征:去掉batch维度,从[1,210,6]转为[210,6] node_features = node_features_tensor.squeeze(0) # 前向传播:传入节点特征、边索引、边特征 predictions = model(node_features, edge_indexes_tensor, edge_features_tensor) # 损失计算 loss_fn = nn.BCELoss() labels_tensor = labels_tensor.float() loss = loss_fn(predictions, labels_tensor) # 反向传播 loss.backward() # optimizer.step() # 训练时启用参数更新
关键修改点
- 调整输入维度定义:将GCNConv的输入维度改为节点特征的6维,而非边特征的4维。
- 修正节点特征形状:把带batch维度的节点特征
[1,210,6]压缩为[210,6],适配GCNConv的输入要求。 - 构建边表示逻辑:通过卷积后的节点特征,提取每条边两端的节点特征,拼接边自身特征后,用全连接层做分类。
- 明确张量作用:GCN的核心是学习节点的上下文特征,边分类需要基于这些节点特征,结合边自身属性完成预测。
内容的提问来源于stack exchange,提问作者Marie-Lyne Roustom
相关产品推荐
相关产品推荐

