基于神经网络输出反推其输入的方法探讨
Got it, let's walk through how to tackle this problem with your activation-function-free neural network, step by step.
First, let's ground ourselves in the setup: we're dealing with a neural network that has no activation functions (so all transformations are strictly linear), we know the biases, weights, and final output, and we need to reverse-engineer the hidden layer's values.
Step 1: Isolate the Weighted Hidden Layer Sum
Since there's no activation function, the final network output is just a linear combination of the hidden layer outputs plus the output layer bias. The first critical step is to strip out that bias to get the pure weighted sum of the hidden layer:
Let’s define our variables clearly to avoid confusion:
O= final network outputb_out= output layer biasW_hidden→out= weight matrix connecting hidden layer to output layerH= hidden layer values we want to solve forThe core equation we start with is:
O = W_hidden→out · H + b_outSubtract the output bias to isolate the weighted hidden sum:
O - b_out = W_hidden→out · H
Step 2: Solve for Hidden Layer Values H
Now we need to use the de-biased output (O - b_out) and our weight matrix to find H. The approach here depends entirely on the structure of the weight matrix:
- If the weight matrix is square and invertible, you can directly compute its inverse and multiply it with the de-biased output to get an exact solution:
H = (W_hidden→out)⁻¹ · (O - b_out) - If the weight matrix isn't square or isn't invertible, things get more nuanced:
- If the system is overdetermined (more output equations than hidden layer nodes), you'll need to use a method like least squares approximation to find the
Hthat best fits the equation (since there's no exact solution). - If it's underdetermined (fewer output equations than hidden layer nodes), there will be infinitely many valid
Hvalues—you'll need additional constraints (like regularization) to narrow down a meaningful solution.
- If the system is overdetermined (more output equations than hidden layer nodes), you'll need to use a method like least squares approximation to find the
Why Substitution Might Not Be Working
You mentioned substitution is a theoretical option but isn't working right now. Here are the most common reasons:
- If the system of equations from
O - b_out = W_hidden→out · His overdetermined, substitution won't give you a single exact solution—you'll end up with conflicting equations. - If it's underdetermined, substitution will leave you with free variables, meaning infinitely many solutions instead of a concrete answer.
- Double-check that you're mapping each hidden node's weight to the correct output term—small indexing mistakes can break substitution entirely.
内容的提问来源于stack exchange,提问作者user8215383

