PyTorch中CNN两种通道实现为何不相等的技术咨询
Great question—this is a super common point of confusion when wrapping your head around CNN channels, so let’s break it down step by step to clear things up.
First, let’s affirm your core intuition: yes, the number of output channels in a Conv2d layer does correspond to the number of independent filters (each with their own weights and bias) applied to the input. But your two implementations differ in critical, easy-to-miss ways that make their outputs unequal:
1. Random Initialization Differences
When you define nn.Conv2d(1, 10, kernel_size=(3, self.embeds_size)), PyTorch initializes 10 distinct filter sets (each of shape (1, 3, self.embeds_size), plus 10 bias values) in one batch.
On the flip side, your ModuleList creates 10 separate nn.Conv2d(1, 1, ...) layers. Each of these layers initializes its own single filter and bias independently. Even though all initializations use the same random distribution, the specific weight values for each filter will be different between the two approaches. Since CNN outputs depend directly on these weights, the initial outputs will never match unless you manually copy the weights over.
2. Tensor Dimension Mismatch
Looking at your post-processing code:
conv2s = torch.stack(conv2s, 1).squeeze(2)
Let’s map the dimensions to spot the problem. Suppose your input x has shape (batch_size, 1, seq_length, embeds_size):
self.conv1(x)outputs a tensor of shape(batch_size, 10, seq_length - 2, 1)(since kernel height is 3, the height dimension shrinks by 2; kernel width matches embeds_size, so width becomes 1).- Each
conv(x)inself.conv2outputs(batch_size, 1, seq_length - 2, 1). Stacking them along dimension 1 gives(batch_size, 10, seq_length - 2, 1). But then yousqueeze(2), which removes theseq_length - 2dimension—resulting in(batch_size, 10, 1).
This means conv1s and conv2s have entirely different shapes, so torch.equal() will always return False. You should instead squeeze the last dimension (dimension 3) to match conv1s:
conv2s = torch.stack(conv2s, 1).squeeze(3)
3. How to Make Them Equivalent
If you want to verify that the two approaches produce the same output (when weights are identical), you can manually copy the weights from conv1 to each layer in conv2:
# After initializing both convolution layers for idx in range(10): # Copy the idx-th filter from conv1 to conv2[idx] self.conv2[idx].weight.data = self.conv1.weight.data[idx:idx+1].clone() self.conv2[idx].bias.data = self.conv1.bias.data[idx:idx+1].clone()
Then use the corrected squeezing step. Now torch.equal(conv1s, conv2s) should return True.
To recap: while both approaches use 10 filters, the independent initialization of the ModuleList layers and incorrect dimension squeezing are why your outputs don’t match. Once you align the weights and fix the tensor shapes, they’ll behave identically.
内容的提问来源于stack exchange,提问作者Matt

