TensorFlow中三个神经网络按序执行操作的问题求助
Hey there! Let's fix your TensorFlow workflow step by step—your core issue is that the order of operations in your loop is completely reversed from what you intended, plus there are a few small issues with how you're handling tensor assignments and mask updates.
First, Let's Clarify the Correct Workflow Order
You want each loop iteration to follow this sequence strictly:
- Update C's weights using A and B's weights (Step 1)
- Train C with the updated weights (Step 2)
- Update A and B's weights using the trained C's weights (Step 3)
- Generate a new mask array (Step 4)
What's Wrong With Your Current Code?
- Your loop runs
training_stepfirst, then updates A/B, then tries to update the mask, then updates C—this is the opposite of your intended flow. - You pre-defined
assanda1outside the loop, but your masktrans_arr1changes every iteration. Those pre-defined assignment operations will use the original mask value, not the updated one. - Updating
trans_arr1withsess.run([trans_arr1])doesn't actually change the tensor that subsequent operations use—you need to properly re-generate or assign the new mask.
Fixed Code Example
Assuming you're using TensorFlow 1.x (since you're using tf.Session), here's the corrected code:
# Make sure all your variables (W1a, W1b, W1c) are initialized first with tf.Session() as sess: sess.run(tf.global_variables_initializer()) # Initialize your mask first trans_arr1 = tf.convert_to_tensor(genRandMat(784, L, ptrans), dtype=tf.float32) for i in range(401): # 400 + 1 iterations as you had # Step 1: Update C's weights using A and B, with current mask update_c_op = W1c.assign(W1a * trans_arr1 + W1b * (tf.ones([784, L]) - trans_arr1)) sess.run(update_c_op) # Step 2: Train C with the freshly updated weights training_step(i, i % 100 == 0, i % 20 == 0) # Step 3: Update A and B's weights using trained C, with current mask # Note: I added B's update since your original code only updated A—adjust if needed update_a_op = W1a.assign(W1c * trans_arr1 + W1a * (tf.ones([784, L]) - trans_arr1)) update_b_op = W1b.assign(W1c * (tf.ones([784, L]) - trans_arr1) + W1b * trans_arr1) sess.run([update_a_op, update_b_op]) # Step 4: Generate new mask and update trans_arr1 # If genRandMat returns a numpy array, convert it to a TF tensor new_mask = tf.convert_to_tensor(genRandMat(784, L, ptrans), dtype=tf.float32) trans_arr1 = new_mask # If trans_arr1 is a TF Variable instead of a tensor, use assign: # trans_arr1.assign(tf.convert_to_tensor(genRandMat(784, L, ptrans), dtype=tf.float32)) # sess.run(trans_arr1) # Trigger the assignment
Key Notes for TensorFlow Beginners
- Strict Order Matters: TensorFlow executes operations in the order you run them with
sess.run()—so always follow your intended sequence exactly. - Dynamic Assignments: Since your mask changes every loop, re-create the assignment operations (
update_c_op,update_a_op, etc.) inside the loop to use the latest mask value. - Mask Handling: If
trans_arr1is a TensorFlowVariable(recommended for mutable state), useassign()to update it instead of reassigning the tensor directly. - Check
training_step: Ensure yourtraining_stepfunction uses the currentW1cweights—if it's using a cached version, your training won't reflect the Step 1 updates.
If you're using TensorFlow 2.x, I'd recommend switching to eager execution (default in TF2) which makes this kind of sequential code much easier to write and debug—you can skip the sess.run() calls entirely and just execute operations line by line.
内容的提问来源于stack exchange,提问作者DarkLord3

