You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于神经网络的(7,4)码编码器:如何实现最小距离3?

How to Adjust Neural Network Architecture to Achieve Minimum Distance 3 for a (7,4) Encoder When Maximizing Minimum Pairwise Distance?

Problem Description

I'm implementing a neural network encoder for a (7,4) code. The theoretical upper bound of the minimum distance for this code is 3, which is achieved by Hamming codes. Essentially, I'm trying to make the model converge to a Hamming code (or a permutation of it that maintains a minimum distance of 3).

Here's my encoder code in TensorFlow:

import numpy as np
import tensorflow as tf

def encoder(message):
    weight1 = np.random.normal(loc=0.0, scale=0.01, size=[16, 32])
    init1 = tf.constant_initializer(weight1)
    out1 = tf.layers.dense(inputs=message, units=32, activation=None, kernel_initializer=init1)
    weight2 = np.random.normal(loc=0.0, scale=0.01, size=[32, 16])
    init2 = tf.constant_initializer(weight2)
    out2 = tf.layers.dense(inputs=out1, units=16, activation=tf.nn.relu, kernel_initializer=init2)
    weight3 = np.random.normal(loc=0.0, scale=0.01, size=[16, 7])
    init3 = tf.constant_initializer(weight3)
    out3 = tf.layers.dense(inputs=out2, units=7, activation=tf.nn.sigmoid, kernel_initializer=init3)
    return out3

During training, to verify if the encoder can reach a minimum distance of 3, I used ground truth data: feeding one-hot vectors to the network and calculating output error against corresponding Hamming code vectors. With the following loss function, the network converges to the Hamming code and achieves a minimum distance of 3:

codeword = encoder(input_bits)
loss = tf.reduce_mean(tf.multiply(0.5, tf.pow(tf.subtract(codeword, hamming_codeword), 2)))

After verification, I changed the loss function to let the encoder find a code with the largest possible minimum distance:

codeword = encoder(input_bits)
distances = []
for i in range(0, 16):
    for j in range(i+1, 16):
        distances.append(tf.linalg.norm(codeword[i] - codeword[j]))
min_dist = tf.reduce_min(distances)
loss = - min_dist

However, the model only reaches a minimum distance close to 2. How should I modify the architecture to achieve the expected minimum distance of 3?


Answer

This is a really interesting problem—switching from "mimicking a known optimal solution" to "independently finding the optimal solution" makes the optimization target much harder, mainly due to the properties of your loss function and the network's expressive capacity. Here are some concrete adjustments you can try:

1. Refactor the Loss Function to Improve Gradient Signals

Your current loss takes the minimum of all pairwise distances and negates it. The problem here is that the gradient is extremely sparse: only the pair of codewords that forms the current minimum distance contributes to parameter updates, while all other pairs have zero gradient. This leads to slow convergence and easily getting stuck in local optima.

Instead, use a margin-based loss that penalizes all pairs whose distance falls below your target margin (note: since your sigmoid outputs are in the 0-1 range, the L2 distance between two distinct binary Hamming codewords is √3 ≈ 1.732, so adjust the margin accordingly):

target_margin = np.sqrt(3)  # Corresponding to L2 distance of binary Hamming code pairs
codeword = encoder(input_bits)

# Generate all pairwise index pairs
i, j = tf.meshgrid(tf.range(16), tf.range(16), indexing='ij')
mask = tf.cast(i < j, tf.float32)  # Avoid duplicate pairs and self-comparisons

# Calculate squared distances to avoid sqrt computation and stabilize gradients
dist_sq = tf.reduce_sum(tf.square(codeword[i] - codeword[j]), axis=-1)

# Loss: penalize pairs where squared distance is less than target margin squared
loss = tf.reduce_mean(mask * tf.maximum(target_margin**2 - dist_sq, 0.0))

This way, every pair with a distance below the target contributes to the gradient, pushing the network to pull those pairs apart—this is far more stable than only optimizing the minimum distance.

2. Adjust Network Architecture for Expressiveness and Output Constraints

Your current architecture uses a sigmoid activation on the output, producing continuous values between 0 and 1. But Hamming codes are strictly binary (0 or 1). Continuous codewords will naturally tend to use intermediate values to meet distance requirements instead of jumping to discrete 0/1 points.

Try these two tweaks:

  • Replace output activation with a straight-through sigmoid: Direct rounding would cause gradient vanishing, so use a straight-through estimator—forward pass uses discrete rounded values, backward pass uses the sigmoid gradient:

    def straight_through_sigmoid(x):
        sigmoid_out = tf.nn.sigmoid(x)
        rounded = tf.round(sigmoid_out)
        # Pass sigmoid gradient during backprop, use rounded value during forward prop
        return rounded + tf.stop_gradient(sigmoid_out - rounded)
    

    Then replace the final sigmoid activation with this function to force outputs close to 0 or 1, aligning with Hamming code properties.

  • Increase network capacity and add batch normalization: Your current 3-layer network (16→32→16→7) might not have enough capacity to find the optimal code without ground truth guidance. Try expanding hidden units or adding batch normalization to stabilize training:

    def encoder(message):
        out1 = tf.layers.dense(inputs=message, units=64, activation=None, 
                              kernel_initializer=tf.random_normal_initializer(stddev=0.01))
        out1 = tf.layers.batch_normalization(out1, training=True)
        out1 = tf.nn.relu(out1)
        
        out2 = tf.layers.dense(inputs=out1, units=32, activation=None, 
                              kernel_initializer=tf.random_normal_initializer(stddev=0.01))
        out2 = tf.layers.batch_normalization(out2, training=True)
        out2 = tf.nn.relu(out2)
        
        out3 = tf.layers.dense(inputs=out2, units=7, activation=None, 
                              kernel_initializer=tf.random_normal_initializer(stddev=0.01))
        return straight_through_sigmoid(out3)
    

    Batch normalization helps the network converge faster and avoids getting trapped in local minima.

3. Optimize Training Strategy

  • Learning rate scheduling: Start with a larger learning rate (e.g., 1e-3) and decay it gradually (e.g., multiply by 0.9 every 1000 epochs) to prevent late-stage training oscillations.
  • Pre-training with ground truth: First train the network using the MSE loss with Hamming code ground truth for a few epochs, then switch to the margin-based loss. Starting from a near-optimal point makes it much easier for the network to converge to the target minimum distance.

Keep in mind that maximizing minimum pairwise distance is a non-convex problem, so combining pre-training and margin-based loss is a reliable way to avoid local optima.

内容的提问来源于stack exchange,提问作者learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:15:10