You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch 0.4.0中求解[6,4,3,1]结构神经网络输出对参数的梯度

Computing Weight Jacobian in PyTorch 0.4.0 for a [6,4,3,1] Network

Got it, let's break this down step by step. You want to calculate the Weight Jacobian (gradient of the network's output with respect to all weights and biases) using Python 3 and PyTorch 0.4.0, with a 4-layer network [6,4,3,1] and 20 input samples. First, let's confirm the parameter count matches your 47 number:

  • Layer 1 (6→4): 6*4 weights + 4 biases = 28 parameters
  • Layer 2 (4→3): 4*3 weights + 3 biases = 15 parameters
  • Layer 3 (3→1): 3*1 weights + 1 bias = 4 parameters
  • Total: 28+15+4=47. Perfect, we're aligned here.

Step 1: Define the Neural Network

First, let's build the network structure using PyTorch's nn.Module:

import torch
import torch.nn as nn

class Net(nn.Module):
    def __init__(self):
        super(Net, self).__init__()
        self.layers = nn.Sequential(
            nn.Linear(6, 4),   # Input layer to hidden layer 1
            nn.ReLU(),         # Activation
            nn.Linear(4, 3),   # Hidden layer 1 to hidden layer 2
            nn.ReLU(),         # Activation
            nn.Linear(3, 1)    # Hidden layer 2 to output layer
        )
    
    def forward(self, x):
        return self.layers(x)

Step 2: Initialize Network and Input Data

Create the network instance and generate your 20 input samples (each with 6 features):

# Initialize the network
net = Net()

# Generate 20 random input samples (shape: [20, 6])
x = torch.randn(20, 6)

Step 3: Compute the Weight Jacobian

The Weight Jacobian here will be a [20, 47] tensor—each row corresponds to the gradient of one sample's output with respect to all 47 parameters. Since PyTorch 0.4.0 doesn't have the higher-level torch.autograd.functional.jacobian (added in later versions), we'll compute this by iterating over each sample and collecting gradients:

# Get all trainable parameters from the network
params = list(net.parameters())

# Initialize Jacobian matrix: [20 samples × 47 parameters]
jacobian = torch.zeros(20, sum(p.numel() for p in params))

for i in range(20):
    # Forward pass for the i-th sample (keep it as a batch to match network input shape)
    sample_output = net(x[i:i+1])  # Shape: [1, 1]
    
    # Compute gradients of the sample's output w.r.t. all parameters
    # Use retain_graph=True to keep the computation graph intact for subsequent iterations
    grads = torch.autograd.grad(sample_output, params, retain_graph=True)
    
    # Flatten all gradients into a single 1D tensor and store in the Jacobian
    flat_grads = torch.cat([grad.view(-1) for grad in grads])
    jacobian[i] = flat_grads

Key Explanations

  • retain_graph=True: This is critical because we need to compute gradients 20 times (once per sample). Without this, PyTorch would destroy the computation graph after the first gradient calculation, causing errors in subsequent iterations.
  • Flattening Parameters/Gradients: Each layer's weights and biases are separate tensors, so we flatten and concatenate them to get a single 47-element vector per sample's gradient—this makes the Jacobian matrix clean and easy to work with.
  • Per-Sample Gradients: If we computed gradients using the entire batch's output directly, PyTorch would return the gradient of the sum of all outputs with respect to parameters. Iterating over samples ensures we get individual gradients for each input-output pair.

Verify Parameter Count

Double-check that we're indeed working with 47 parameters:

total_params = sum(p.numel() for p in net.parameters())
print(f"Total trainable parameters: {total_params}")  # Should print 47

Optional: Batch-Summed Gradient

If you only need the gradient of the summed batch output (instead of per-sample Jacobian), you can skip the loop and do this:

batch_output = net(x).sum()
batch_grads = torch.autograd.grad(batch_output, params)
flat_batch_grads = torch.cat([g.view(-1) for g in batch_grads])

内容的提问来源于stack exchange,提问作者Sibghat Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:07:26