如何使用不可微损失函数?神经网络码本优化梯度问题求助
I'm trying to generate a codebook at the output of a fully connected neural network, where the goal is to maximize the minimum Euclidean distance between any pair of codebook points. For example, when input dimension is 2 and output dimension is 3, the optimal mapping (up to permutation) is:
- 00 → 000
- 01 → 011
- 10 → 101
- 11 → 110
I wrote the following TensorFlow code, but I'm hitting a ValueError: No gradients provided for any variable because the loss function seems non-differentiable. Here's my code:
import tensorflow as tf import numpy as np import itertools input_bits = tf.placeholder(dtype=tf.float32, shape=[None, 2], name='input_bits') code_out = tf.placeholder(dtype=tf.float32, shape=[None, 3], name='code_out') np.random.seed(1331) def find_code(message): weight1 = np.random.normal(loc=0.0, scale=0.01, size=[2, 3]) init1 = tf.constant_initializer(weight1) out = tf.layers.dense(inputs=message, units=3, activation=tf.nn.sigmoid, kernel_initializer=init1) return out code = find_code(input_bits) distances = [] for i in range(0, 3): for j in range(i+1, 3): distances.append(tf.linalg.norm(code_out[i]-code_out[j])) min_dist = tf.reduce_min(distances) # avg_dist = tf.reduce_mean(distances) loss = -min_dist opt = tf.train.AdamOptimizer().minimize(loss) init_variables = tf.global_variables_initializer() sess = tf.Session() sess.run(init_variables) saver = tf.train.Saver() count = int(1e4) for i in range(count): input_bit = [list(k) for k in itertools.product([0, 1], repeat=2)] code_preview = sess.run(code, feed_dict={input_bits: input_bit}) sess.run(opt, feed_dict={input_bits: input_bit, code_out: code_preview})
Why You're Getting the Gradient Error
The core issue is two-fold:
- Disconnected computational graph: You're passing
code_preview(a static tensor from a previous session run) into thecode_outplaceholder. This tensor has no link to the network's weights, so TensorFlow can't compute gradients from the loss back to your model parameters. - Non-differentiable
minoperation: Even if you fixed the graph link,tf.reduce_minisn't fully differentiable (its gradient is undefined when multiple distances equal the minimum), which breaks gradient-based optimizers like Adam.
Fixes and Alternative Approaches
Here's how to resolve both problems and train your network effectively:
1. Keep the Graph Connected
First, remove the code_out placeholder entirely. Compute pairwise distances directly using the code tensor (the network's live output) so gradients can flow from the loss back to the weights.
2. Use a Smooth Surrogate for the Minimum Distance
Since the raw minimum is non-differentiable, replace it with a smooth approximation that still encourages maximizing the smallest pairwise distance. Here are two reliable options:
Option A: Log-Sum-Exp (LSE) Smoothed Minimum
This trick approximates the minimum using exponential functions, creating a differentiable surrogate. Adjust the beta parameter to control how close it gets to the true minimum (higher beta = closer to true min, but risk of numerical instability):
# Compute all pairwise distances from the network's output pairwise_dists = [] num_codes = tf.shape(code)[0] for i in range(num_codes): for j in range(i+1, num_codes): dist = tf.linalg.norm(code[i] - code[j]) pairwise_dists.append(dist) pairwise_dists = tf.stack(pairwise_dists) # Smooth min approximation beta = 5.0 # Start with 1.0, increase gradually if training is stable smoothed_min = -tf.math.log(tf.reduce_sum(tf.math.exp(-beta * pairwise_dists))) / beta loss = -smoothed_min # Maximize the smoothed minimum distance
Option B: Margin-Based Loss
Instead of maximizing the minimum distance, enforce that all pairs are at least a target margin apart. This loss penalizes pairs that fall below the margin and is fully differentiable:
margin = 1.0 # Adjust based on your desired minimum distance pairwise_dists = [] num_codes = tf.shape(code)[0] for i in range(num_codes): for j in range(i+1, num_codes): dist = tf.linalg.norm(code[i] - code[j]) # Penalize distances below the margin pairwise_dists.append(tf.maximum(margin - dist, 0.0)) loss = tf.reduce_sum(pairwise_dists)
Revised Full Code Example
Here's your code updated with the LSE smoothed minimum approach:
import tensorflow as tf import numpy as np import itertools input_bits = tf.placeholder(dtype=tf.float32, shape=[None, 2], name='input_bits') np.random.seed(1331) def find_code(message): weight1 = np.random.normal(loc=0.0, scale=0.01, size=[2, 3]) init1 = tf.constant_initializer(weight1) out = tf.layers.dense(inputs=message, units=3, activation=tf.nn.sigmoid, kernel_initializer=init1) return out code = find_code(input_bits) # Compute all pairwise distances pairwise_dists = [] num_codes = tf.shape(code)[0] for i in range(num_codes): for j in range(i+1, num_codes): dist = tf.linalg.norm(code[i] - code[j]) pairwise_dists.append(dist) pairwise_dists = tf.stack(pairwise_dists) # Smooth min approximation beta = 5.0 smoothed_min = -tf.math.log(tf.reduce_sum(tf.math.exp(-beta * pairwise_dists))) / beta loss = -smoothed_min opt = tf.train.AdamOptimizer(learning_rate=0.01).minimize(loss) init_variables = tf.global_variables_initializer() sess = tf.Session() sess.run(init_variables) saver = tf.train.Saver() count = int(1e4) # Precompute input bits once (no need to regenerate every iteration) input_bit = [list(k) for k in itertools.product([0, 1], repeat=2)] for i in range(count): _, current_loss, current_codes = sess.run([opt, loss, code], feed_dict={input_bits: input_bit}) # Print progress every 1000 steps if i % 1000 == 0: print(f"Step {i}, Loss: {current_loss:.4f}") print("Current codebook:") print(np.round(current_codes, 3)) print("---")
Key Notes
- Avoid disconnected tensors: Always compute loss directly from the network's output to maintain gradient flow.
- Tune hyperparameters: Adjust
beta(for LSE) ormargin(for margin loss) based on training stability and your desired codebook properties. - Monitor progress: Print the codebook periodically to verify that the minimum distance is increasing over time.
内容的提问来源于stack exchange,提问作者learner

