TensorFlow MNIST指南中矩阵运算代码的理解困惑
Hey Andrew, let's clear up this confusion step by step—you're actually closer to the right understanding than you might think!
First, let's recap the shapes of the tensors we're working with:
x: Has a shape of[None, 784]. TheNonemeans it can hold any number of MNIST images, and each image is flattened into a 784-dimensional vector (one value per pixel in the 28x28 image). So if you feed in 50 images,xbecomes a[50, 784]matrix.W: Has a shape of[784, 10]. Think of this as 10 separate 784-dimensional weight vectors (one for each digit class, 0 through 9). Each column inWcorresponds to the weights for a single digit class.
Breaking Down tf.matmul(x, W)
When you multiply x and W, here's what happens at a granular level:
For a single image (a single row in x, shape [784]), multiplying by W gives you a 10-dimensional vector. Each element in this vector is the sum of multiplying every pixel value in the image by its corresponding weight for that digit class.
For example, the 4th element in the resulting vector (index 3, since we start counting at 0) is:x₁*W₁₄ + x₂*W₂₄ + ... + x₇₈₄*W₇₈₄₄
This is exactly the extended version of the simplified example you saw ((W₁₁*x₁)+(W₁₂*x₂)+(W₁₃*x₃)—the example just uses 3 pixels instead of 784, and refers to a single digit class's weights).
Your initial thought was correct: this operation does sum the product of each pixel's value with the corresponding weight for each digit class. The confusion might have come from how the example's notation maps to the tensor shapes.
Batch Processing Bonus
Since x can hold multiple images, tf.matmul(x, W) does this calculation for all images at once. The output shape becomes [None, 10], where each row is the 10 raw "scores" for one image (one score per digit class). Adding the bias b (which gets broadcast to every row) and applying softmax then converts those scores into probabilities that sum to 1 for each image.
To make this even clearer, here's a tiny simplified example:
If x was [[1,2,3], [4,5,6]] (2 images, 3 pixels each) and W was [[0.1, 0.4], [0.2, 0.5], [0.3, 0.6]] (3 pixels, 2 classes), the matrix multiplication would result in:
[ [1*0.1 + 2*0.2 + 3*0.3, 1*0.4 + 2*0.5 + 3*0.6], [4*0.1 + 5*0.2 + 6*0.3, 4*0.4 + 5*0.5 + 6*0.6] ]
Each row's values are the summed pixel-weight products for each class—exactly what you expected!
So to wrap up: your understanding isn't wrong at all. The example's simplified notation just hides the full dimension of the problem, but the core logic of summing pixel-weight products per class is exactly what's happening here.
内容的提问来源于stack exchange,提问作者Andrew

