You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用不可微损失函数?神经网络码本优化梯度问题求助

How to Maximize the Minimum Euclidean Distance Between Neural Network Outputs for Codebook Generation?

I'm trying to generate a codebook at the output of a fully connected neural network, where the goal is to maximize the minimum Euclidean distance between any pair of codebook points. For example, when input dimension is 2 and output dimension is 3, the optimal mapping (up to permutation) is:

  • 00 → 000
  • 01 → 011
  • 10 → 101
  • 11 → 110

I wrote the following TensorFlow code, but I'm hitting a ValueError: No gradients provided for any variable because the loss function seems non-differentiable. Here's my code:

import tensorflow as tf
import numpy as np
import itertools

input_bits = tf.placeholder(dtype=tf.float32, shape=[None, 2], name='input_bits')
code_out = tf.placeholder(dtype=tf.float32, shape=[None, 3], name='code_out')

np.random.seed(1331)

def find_code(message):
    weight1 = np.random.normal(loc=0.0, scale=0.01, size=[2, 3])
    init1 = tf.constant_initializer(weight1)
    out = tf.layers.dense(inputs=message, units=3, activation=tf.nn.sigmoid, kernel_initializer=init1)
    return out

code = find_code(input_bits)

distances = []
for i in range(0, 3):
    for j in range(i+1, 3):
        distances.append(tf.linalg.norm(code_out[i]-code_out[j]))

min_dist = tf.reduce_min(distances)
# avg_dist = tf.reduce_mean(distances)
loss = -min_dist

opt = tf.train.AdamOptimizer().minimize(loss)

init_variables = tf.global_variables_initializer()
sess = tf.Session()
sess.run(init_variables)
saver = tf.train.Saver()

count = int(1e4)
for i in range(count):
    input_bit = [list(k) for k in itertools.product([0, 1], repeat=2)]
    code_preview = sess.run(code, feed_dict={input_bits: input_bit})
    sess.run(opt, feed_dict={input_bits: input_bit, code_out: code_preview})

Why You're Getting the Gradient Error

The core issue is two-fold:

  1. Disconnected computational graph: You're passing code_preview (a static tensor from a previous session run) into the code_out placeholder. This tensor has no link to the network's weights, so TensorFlow can't compute gradients from the loss back to your model parameters.
  2. Non-differentiable min operation: Even if you fixed the graph link, tf.reduce_min isn't fully differentiable (its gradient is undefined when multiple distances equal the minimum), which breaks gradient-based optimizers like Adam.

Fixes and Alternative Approaches

Here's how to resolve both problems and train your network effectively:

1. Keep the Graph Connected

First, remove the code_out placeholder entirely. Compute pairwise distances directly using the code tensor (the network's live output) so gradients can flow from the loss back to the weights.

2. Use a Smooth Surrogate for the Minimum Distance

Since the raw minimum is non-differentiable, replace it with a smooth approximation that still encourages maximizing the smallest pairwise distance. Here are two reliable options:

Option A: Log-Sum-Exp (LSE) Smoothed Minimum

This trick approximates the minimum using exponential functions, creating a differentiable surrogate. Adjust the beta parameter to control how close it gets to the true minimum (higher beta = closer to true min, but risk of numerical instability):

# Compute all pairwise distances from the network's output
pairwise_dists = []
num_codes = tf.shape(code)[0]
for i in range(num_codes):
    for j in range(i+1, num_codes):
        dist = tf.linalg.norm(code[i] - code[j])
        pairwise_dists.append(dist)
pairwise_dists = tf.stack(pairwise_dists)

# Smooth min approximation
beta = 5.0  # Start with 1.0, increase gradually if training is stable
smoothed_min = -tf.math.log(tf.reduce_sum(tf.math.exp(-beta * pairwise_dists))) / beta
loss = -smoothed_min  # Maximize the smoothed minimum distance
Option B: Margin-Based Loss

Instead of maximizing the minimum distance, enforce that all pairs are at least a target margin apart. This loss penalizes pairs that fall below the margin and is fully differentiable:

margin = 1.0  # Adjust based on your desired minimum distance
pairwise_dists = []
num_codes = tf.shape(code)[0]
for i in range(num_codes):
    for j in range(i+1, num_codes):
        dist = tf.linalg.norm(code[i] - code[j])
        # Penalize distances below the margin
        pairwise_dists.append(tf.maximum(margin - dist, 0.0))
loss = tf.reduce_sum(pairwise_dists)

Revised Full Code Example

Here's your code updated with the LSE smoothed minimum approach:

import tensorflow as tf
import numpy as np
import itertools

input_bits = tf.placeholder(dtype=tf.float32, shape=[None, 2], name='input_bits')

np.random.seed(1331)

def find_code(message):
    weight1 = np.random.normal(loc=0.0, scale=0.01, size=[2, 3])
    init1 = tf.constant_initializer(weight1)
    out = tf.layers.dense(inputs=message, units=3, activation=tf.nn.sigmoid, kernel_initializer=init1)
    return out

code = find_code(input_bits)

# Compute all pairwise distances
pairwise_dists = []
num_codes = tf.shape(code)[0]
for i in range(num_codes):
    for j in range(i+1, num_codes):
        dist = tf.linalg.norm(code[i] - code[j])
        pairwise_dists.append(dist)
pairwise_dists = tf.stack(pairwise_dists)

# Smooth min approximation
beta = 5.0
smoothed_min = -tf.math.log(tf.reduce_sum(tf.math.exp(-beta * pairwise_dists))) / beta
loss = -smoothed_min

opt = tf.train.AdamOptimizer(learning_rate=0.01).minimize(loss)

init_variables = tf.global_variables_initializer()
sess = tf.Session()
sess.run(init_variables)
saver = tf.train.Saver()

count = int(1e4)
# Precompute input bits once (no need to regenerate every iteration)
input_bit = [list(k) for k in itertools.product([0, 1], repeat=2)]
for i in range(count):
    _, current_loss, current_codes = sess.run([opt, loss, code], feed_dict={input_bits: input_bit})
    # Print progress every 1000 steps
    if i % 1000 == 0:
        print(f"Step {i}, Loss: {current_loss:.4f}")
        print("Current codebook:")
        print(np.round(current_codes, 3))
        print("---")

Key Notes

  • Avoid disconnected tensors: Always compute loss directly from the network's output to maintain gradient flow.
  • Tune hyperparameters: Adjust beta (for LSE) or margin (for margin loss) based on training stability and your desired codebook properties.
  • Monitor progress: Print the codebook periodically to verify that the minimum distance is increasing over time.

内容的提问来源于stack exchange,提问作者learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:16:55