PyTorch中稠密-稀疏矩阵反向传播报错问题求助
自定义稀疏矩阵广播Autograd类的反向传播错误问题
PyTorch目前暂不支持稀疏矩阵的广播功能,我实现了一个简单的autograd类,前向传播运行正常,但反向传播时出现如下错误:
Could not run 'aten::as_strided' with arguments from the 'SparseCPU' backend.
复现该问题的最小代码片段:
import torch from torch.autograd import Function # 创建这个自定义autograd函数是为了处理稀疏矩阵的广播 class BroadcastSparse1D(Function): @staticmethod def forward(ctx, sparse_1d, target_shape): assert sparse_1d.dim() == 1, "输入张量必须是一维的" target_dims = len(target_shape) new_indices = torch.zeros(target_dims, sparse_1d._nnz()) new_indices[0, :] = sparse_1d._indices()[0, :] broadcasted_sparse = torch.sparse_coo_tensor(new_indices, sparse_1d._values(), target_shape, device=sparse_1d.device) ctx.save_for_backward(sparse_1d) ctx.target_shape = target_shape return broadcasted_sparse @staticmethod def backward(ctx, grad_output): sparse_1d, = ctx.saved_tensors # 将grad_output转为稠密张量并对指定维度求和 grad_output_dense = grad_output.to_dense() grad_sum_dense = grad_output_dense.sum(dim=tuple(range(1, len(ctx.target_shape)))) # 将求和后的稠密梯度转回稀疏张量 grad_sum_sparse = grad_sum_dense.to_sparse() grad_input = torch.sparse_coo_tensor(sparse_1d._indices(), grad_sum_sparse._values(), sparse_1d.shape, device=grad_output.device) return grad_input, None broadcast_sparse_1d = BroadcastSparse1D.apply # 构造输入参数: # 我在跟踪`a`和`c`的梯度,另外`c`和`d`是稀疏张量 a = torch.rand((4, 4, 176)) a.requires_grad_(True) b = torch.rand((64, 64)) c = torch.rand((64, 64, 184)).to_sparse() c.requires_grad_(True) d = torch.rand((184, 4, 4, 176)).to_sparse() # 前向计算流程 e = torch.sum(b.unsqueeze(-1) * c, dim=(0, 1)) e_broadcasted = broadcast_sparse_1d(e, d.shape) # 用这个模拟广播,`e`形状为(184),`d`形状为(184, 4, 4, 176) f = torch.sum(e_broadcasted * d, dim=0) result = -torch.min((a * 400) - f)
请问我是否犯了什么明显的错误?
内容的提问来源于stack exchange,提问作者mino
相关产品推荐
相关产品推荐

