You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

图卷积网络(GCN)中Antithetic Sampling方差缩减实现问题排查

对偶采样(Antithetic Sampling)用于图顶点采样的实现问题

我正尝试实现**对偶采样(Antithetic Sampling)**来对图的顶点进行采样,并在采样后的图上训练下游图卷积网络(GCN)模型。

对偶采样是一种方差缩减技术,通过生成成对的随机样本及其对应的对偶样本抵消波动,从而得到更准确的估计结果。

以下是我的代码,核心逻辑在_Antithetic_sampling(...)函数中:

import math
import torch
import numpy as np
import scipy.sparse as sp

from scipy.sparse.linalg import norm as sparse_norm
class Antithetic_Sampler(Sampler):
    def __init__(self, pre_probs, features, adj, **kwargs):
        super().__init__(features, adj, **kwargs)
        col_norm = sparse_norm(adj, axis=0)
        self.probs = col_norm / np.sum(col_norm)
    def sampling(self, v):
        """
        Inputs:
            v: batch nodes list
        """
        all_support = [[]] * self.num_layers  # Initialize empty list for all layers
        cur_out_nodes = v 
        for layer_index in range(self.num_layers - 1, -1, -1): # Start from the last layer and move backwards
            cur_sampled, cur_support = self._Antithetic_sampling(cur_out_nodes, self.layer_sizes[layer_index])  # sample nodes and collect support
            all_support[layer_index] = cur_support # for corresponding layer, Store current support in all_support 
            cur_out_nodes = cur_sampled # Update nodes 
        all_support = self._change_sparse_to_tensor(all_support) # Convert support to tensor representation       
        sampled_X0 = self.features[cur_out_nodes]  # Extract features of the sampled nodes 
        return sampled_X0, all_support, 0

    # Perform Antithetic Sampling
    def _Antithetic_sampling(self, v_indices, output_size):
        support = self.adj[v_indices, :]
        neis = np.nonzero(np.sum(support, axis=0))[1]
        # Create two sets of random sampling weights
        p1 = self.probs[neis]
        p1 = p1 / np.sum(p1) # Normalize probability 
        p2 = 1 - p1
        p2 = p2 / np.sum(p2)

        # Sample the first set of neighbors
        sampled_1 = np.random.choice(np.arange(np.size(neis)), output_size, True, p1)
        u_sampled_1 = neis[sampled_1]
        support_1 = support[:, u_sampled_1]
        sampled_p1 = p1[sampled_1]
        support_1 = support_1.dot(sp.diags(1.0 / (sampled_p1 * output_size)))

        # Sample the second set of neighbors with opposite weights

        sampled_2 = np.random.choice(np.arange(np.size(neis)), output_size, True, p2)
        u_sampled_2 = neis[sampled_2]
        support_2 = support[:, u_sampled_2]
        sampled_p2 = p2[sampled_2]
        support_2 = support_2.dot(sp.diags(1.0 / (sampled_p2 * output_size)))

        # Average two sets of sampled neighbors
        u_sampled = (u_sampled_1 + u_sampled_2) // 2
        support = (support_1 + support_2) / 2
        #print("U samples: ",  u_sampled)
        #print("Support: ",  support)

        return u_sampled, support

代码运行输出如下:

epchs:0~9 => test_loss: 1.882, test_acc: 0.319
epchs:10~19 => test_loss: 1.879, test_acc: 0.319
epchs:20~29 => test_loss: 1.876, test_acc: 0.319
epchs:30~39 => test_loss: 1.879, test_acc: 0.319
epchs:40~49 => test_loss: 1.871, test_acc: 0.319
epchs:50~59 => test_loss: 1.873, test_acc: 0.319

可以看到测试准确率始终没有变化,说明我的对偶采样实现可能存在逻辑错误,希望能帮忙排查问题。

关于对偶采样的更多细节可查看相关教材的第8.2节。


问题排查与修复建议

1. 核心采样逻辑完全偏离对偶采样定义

你当前实现中u_sampled = (u_sampled_1 + u_sampled_2) // 2这一步是完全错误的:对偶采样的核心是生成负相关的样本对,而非将两个样本的节点索引直接取平均——这种操作会得到大量无效节点(比如非整数索引、不存在的节点),直接导致采样后的节点集合完全失效,模型无法学到有效信息。

正确的做法是保留两组独立的采样节点集合,在后续GCN前向传播中分别计算输出,再将两个输出取平均,以此实现方差缩减。

2. 对偶样本权重生成错误

用p2 = 1 - p1生成对偶采样权重的方式不成立:

  • 当p1接近0或1时,p2会出现极端值,采样结果会严重偏向少数节点
  • 这种方式无法保证两组样本的负相关性,完全失去对偶采样的方差缩减作用

正确的对偶样本生成应该基于逆变换采样:

# 生成原样本的均匀分布采样值
uniform_samples = np.random.uniform(0, 1, output_size)
# 生成对偶样本的均匀分布采样值(与原样本负相关)
antithetic_uniform = 1 - uniform_samples

# 基于累积分布函数(CDF)采样
cdf = np.cumsum(p1)
sampled_1 = np.searchsorted(cdf, uniform_samples)
sampled_2 = np.searchsorted(cdf, antithetic_uniform)

这样生成的两组样本严格负相关,符合对偶采样的核心要求。

3. 支持矩阵与采样节点的处理错误

你当前直接合并两个支持矩阵的操作毫无意义,因为节点集合已经失效。正确的做法是保留两组支持矩阵,在模型前向传播时分别计算特征聚合结果,再取平均。

修复后的核心函数示例

def _Antithetic_sampling(self, v_indices, output_size):
    support = self.adj[v_indices, :]
    neis = np.nonzero(np.sum(support, axis=0))[1]
    p1 = self.probs[neis]
    p1 = p1 / np.sum(p1)  # 归一化概率
    cdf = np.cumsum(p1)

    # 生成原样本和对偶样本的均匀分布采样值
    uniform_samples = np.random.uniform(0, 1, output_size)
    antithetic_uniform = 1 - uniform_samples

    # 基于CDF完成采样
    sampled_1 = np.searchsorted(cdf, uniform_samples)
    sampled_2 = np.searchsorted(cdf, antithetic_uniform)

    u_sampled_1 = neis[sampled_1]
    u_sampled_2 = neis[sampled_2]

    # 构建两组支持矩阵
    support_1 = support[:, u_sampled_1]
    support_1 = support_1.dot(sp.diags(1.0 / (p1[sampled_1] * output_size)))

    support_2 = support[:, u_sampled_2]
    support_2 = support_2.dot(sp.diags(1.0 / (p1[sampled_2] * output_size)))

    # 返回两组采样结果,后续在模型中取平均
    return (u_sampled_1, u_sampled_2), (support_1, support_2)

同时需要修改sampling函数,将两组结果传递给模型,并在GCN的前向传播逻辑中对两组输出取平均,才能真正发挥对偶采样的方差缩减作用。

4. 其他潜在问题

  • 初始化all_support时,[[]] * self.num_layers会导致所有列表元素指向同一个对象,应改为all_support = [[] for _ in range(self.num_layers)]
  • 确认基类Sampler的_change_sparse_to_tensor方法能正确处理两组支持矩阵的格式

内容的提问来源于stack exchange,提问作者Asif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 18:06:07