TensorFlow 1.2.1报错ValueError:无梯度提供,求问题排查解决
Hey, let's dig into why you're hitting that gradient error. After looking through your code, there are two critical issues breaking the backpropagation path, plus a couple of smaller fixes to get everything working smoothly.
1. The big culprit: tf.arg_max is non-differentiable
In your inference function, the line logit = tf.cast(tf.arg_max(logit, 1), tf.float32) completely kills gradient flow. tf.arg_max outputs discrete class indices (0, 1, 2 for your wine dataset), which aren't continuous values. When you pass this to your loss function, TensorFlow can't compute gradients back to your model's weights and biases—there's no way to calculate how changing a weight affects a discrete index.
You need to remove this line entirely. The tf.nn.softmax_cross_entropy_with_logits function expects the raw model outputs (logits, before softmax or argmax) to compute loss correctly.
Also, you were defining two variables named weight and bias in the same variable_scope('layer1')—this would cause a variable reuse error. I moved the second layer into its own scope to avoid that.
Fixed inference function:
def inference(x): with tf.variable_scope('layer1'): weight1 = tf.get_variable('weight', [13, 7], initializer=tf.truncated_normal_initializer(stddev=0.1)) bias1 = tf.get_variable('bias', [7], initializer=tf.constant_initializer(0.1)) layer1 = tf.nn.relu(tf.matmul(x, weight1) + bias1) # Move layer 2 to its own scope to avoid variable name collisions with tf.variable_scope('layer2'): weight2 = tf.get_variable('weight', [7, 3], initializer=tf.truncated_normal_initializer(stddev=0.1)) bias2 = tf.get_variable('bias', [3], initializer=tf.constant_initializer(0.1)) logit = tf.matmul(layer1, weight2) + bias2 return logit
2. Fix the label format for cross-entropy
tf.nn.softmax_cross_entropy_with_logits requires labels to be one-hot encoded, but you're passing raw integer labels (0, 1, 2) as floats. Here's how to fix this:
- Change the
y_placeholder to accept integer values (since we'll convert to one-hot) - Use
tf.one_hotto convert your labels into the required shape
Updated code for the loss calculation:
# Update placeholder to int32 y_ = tf.placeholder(tf.int32, [None]) # Convert labels to one-hot encoding (3 classes for wine dataset) cross_entropy = tf.nn.softmax_cross_entropy_with_logits( labels=tf.one_hot(y_, depth=3), logits=y ) cross_entropy_mean = tf.reduce_mean(cross_entropy)
3. Correct accuracy calculation
Your original accuracy check compared the raw logits (y) directly to the labels (y_), which doesn't make sense. You need to take the argmax of the logits to get the predicted class index, then compare that to the true labels:
correct_prediction = tf.equal(tf.argmax(y, 1), tf.cast(y_, tf.int64)) accuracy = tf.reduce_mean(tf.cast(correct_prediction, tf.float32))
Full working code
Putting it all together, here's the corrected code that should run without the gradient error:
from sklearn import datasets import random import tensorflow as tf wine = datasets.load_wine() def generate_batch(batch_size, wine): batch_x = [] batch_y = [] for _ in range(batch_size): index = random.randint(0, 177) batch_y.append(wine.target[index]) # Keep as integer for one-hot conversion batch_x.append(wine.data[index]) return batch_x, batch_y def inference(x): with tf.variable_scope('layer1'): weight1 = tf.get_variable('weight', [13, 7], initializer=tf.truncated_normal_initializer(stddev=0.1)) bias1 = tf.get_variable('bias', [7], initializer=tf.constant_initializer(0.1)) layer1 = tf.nn.relu(tf.matmul(x, weight1) + bias1) with tf.variable_scope('layer2'): weight2 = tf.get_variable('weight', [7, 3], initializer=tf.truncated_normal_initializer(stddev=0.1)) bias2 = tf.get_variable('bias', [3], initializer=tf.constant_initializer(0.1)) logit = tf.matmul(layer1, weight2) + bias2 return logit x = tf.placeholder(tf.float32, [None, 13]) y_ = tf.placeholder(tf.int32, [None]) y = inference(x) cross_entropy = tf.nn.softmax_cross_entropy_with_logits( labels=tf.one_hot(y_, depth=3), logits=y ) cross_entropy_mean = tf.reduce_mean(cross_entropy) train_step = tf.train.GradientDescentOptimizer(0.001).minimize(cross_entropy_mean) correct_prediction = tf.equal(tf.argmax(y, 1), tf.cast(y_, tf.int64)) accuracy = tf.reduce_mean(tf.cast(correct_prediction, tf.float32)) with tf.Session() as sess: tf.global_variables_initializer().run() for i in range(2000): data, target = generate_batch(20, wine) _, loss, acc = sess.run([train_step, cross_entropy_mean, accuracy], feed_dict={x: data, y_: target}) if i % 200 == 0: print(f"Step {i}, Loss: {loss:.4f}, Accuracy: {acc:.4f}")
To recap: The main issue was using tf.arg_max which breaks gradient flow. Fixing that, plus adjusting the label format and accuracy calculation, gets your model training properly.
内容的提问来源于stack exchange,提问作者li.SQ

