TensorFlow中reduce_sum处理placeholder报错及Replicated Softmax模型实现问题
First, let's tackle the InvalidArgumentError: 'In[0] is not a matrix' from your outer product calculation—this has a clear root cause and fix.
The Problem with Your Outer Product Code
Your line:
tf.matmul(tf.squeeze(tf.reduce_sum(visible_units, 1, keepdims=True)), tf.expand_dims(self.h, 1))
has two critical issues:
tf.reduce_sum(visible_units, 1, keepdims=True)returns a 2D tensor of shape[batch_size, 1]—which is exactly what you need. Buttf.squeeze()strips the singleton dimension, turning it into a 1D tensor[batch_size].tf.matmul()requires both inputs to be at least 2D matrices, hence the "not a matrix" error.- You're expanding
self.halong axis 1, which turns a 1D tensor[n_hidden]into[n_hidden, 1]. Even if you fixed the first issue, multiplying[batch_size,1]with[n_hidden,1]would fail due to incompatible inner dimensions.
Correct Way to Compute the Outer Product
To get the outer product of your row-wise sum (a batch of column vectors) and the hidden bias self.h, you need:
- The sum tensor to stay as
[batch_size, 1](no squeeze!) self.hreshaped to[1, n_hidden]so the inner dimensions match for matrix multiplication
Here's the fixed code:
# Compute row-wise sum, keep it as 2D (batch_size x 1) sum_visible = tf.reduce_sum(visible_units, axis=1, keepdims=True) # Reshape self.h from [n_hidden] to [1, n_hidden] h_reshaped = tf.expand_dims(self.h, axis=0) # Compute outer product: (batch_size x 1) * (1 x n_hidden) = batch_size x n_hidden outer_product = tf.matmul(sum_visible, h_reshaped)
This works because tf.matmul expects matching inner dimensions (1 and 1 here), and outputs the exact batch_size x n_hidden tensor you need for Replicated Softmax.
Troubleshooting the Random reduce_sum Error with tf.placeholder
Since you mentioned a "random error" without specific logs, here are the most likely culprits to investigate:
- Mismatched input dimensions: If your placeholder uses dynamic shapes (e.g.,
shape=[None, n_visible_units]), double-check that every batch of input data has exactly 2 dimensions. A random batch with a missing dimension (e.g.,[batch_size]instead of[batch_size, n_visible_units]) will trigger a reduce_sum failure. - Edge case batch sizes: If some batches have a size of 0 (from a poorly split dataset, for example), reduce_sum might throw unexpected errors. Add checks to validate your input batch sizes before feeding them to the graph.
- Graph/session conflicts: Reusing the same placeholder across multiple graphs or sessions can cause random tensor shape mismatches. Ensure your placeholder is tied to exactly one graph, and you're feeding data consistently to that session.
- Resource contention: In rare cases, random errors can stem from GPU/CPU resource conflicts. Try running your code on CPU temporarily to rule out hardware-related issues.
Always capture the full stack trace when the random error occurs—it will give you specific details about whether it's a shape mismatch, out-of-memory error, or something else.
内容的提问来源于stack exchange,提问作者Avinash

