TensorFlow梯度异常:优化3D源位置时无梯度提供
问题描述
尝试优化多项式系数以估计3D空间中源的位置,2D示例正常运行,但切换到真实正向模型后,initial_poly_coefficients与计算图断开连接,所有系数的梯度返回None,触发报错:
ValueError: No gradients provided for any variable: ['Variable:0', 'Variable:0', 'Variable:0', 'Variable:0', 'Variable:0'].
已做调试尝试:
- 关闭eager execution模式后,TensorFlow可追踪梯度,但失去TensorFlow与NumPy的交互功能;
- 通过
tf.test.compute_gradient验证magnetic_dipole_fast_tf和generate_H均具备可微性; - 怀疑某处赋值操作断开了多项式系数与损失函数的关联。
问题根源
问题出在generate_H函数中:
- 用
np.zeros初始化了NumPy数组H,后续通过H[i, j, :] = lf将TensorFlow张量的值赋值给NumPy数组——这个过程会切断梯度传播链,因为NumPy数组不属于TensorFlow的计算图,无法追踪梯度。 - 最后将NumPy数组转为TensorFlow张量时,该张量已和上游的
poly_coefficients失去梯度关联。
修复方案
修改generate_H函数,全程使用TensorFlow操作,避免NumPy数组介入:
- 用
tf.zeros替代np.zeros初始化张量; - 用
tf.tensor_scatter_nd_update进行索引赋值,替代NumPy的直接赋值,保证梯度正常传播。
完整修复代码
import numpy as np import tensorflow as tf import matplotlib.pyplot as plt from tensorflow.python.ops.numpy_ops import np_config np_config.enable_numpy_behavior() # Simulated data num_sources = 50 num_channels_x = 5 num_channels_y = 5 # Define the ground truth positions of the sources in 3D space (x, y, z coordinates) source_positions_x = tf.cast(np.linspace(-1, 1, num_sources), dtype=tf.float32) source_positions_y = tf.zeros_like(source_positions_x) # All sources at y=0 source_positions_z = tf.cast(-(source_positions_x ** 2) + source_positions_x -1.0, dtype=tf.float32) # Specify the coefficients as constants coefficients_values = [0.01, -1.2, 1.1, 0.0, 2.0] # Create TensorFlow variables with these specific values initial_poly_coefficients = [tf.Variable([value], dtype=tf.float32,trainable=True) for value in coefficients_values] # Define the initial guess positions of the sources initial_source_positions_x = tf.cast(np.linspace(-1, 1, num_sources), dtype=tf.float32) initial_source_positions_y = tf.zeros_like(initial_source_positions_x) # Initial guess: all sources at z=0 initial_source_positions_z = tf.math.polyval(initial_poly_coefficients, initial_source_positions_x) # Define the positions of the channels (on a 2D grid in the xy-plane above the sources) channel_positions_x = np.linspace(-1, 1, num_channels_x) channel_positions_y = np.linspace(-1, 1, num_channels_y) channel_positions_z = tf.zeros((num_channels_x * num_channels_y,), dtype=tf.float32) # All channels at z=0 # Create 3D arrays for positions and orientations (in this example, orientations are along the z-axis) source_positions = tf.stack([source_positions_x, source_positions_y, source_positions_z], axis=1) source_orientations = tf.constant([[0.0, 0.0, 1.0]] * num_sources, dtype=tf.float32) # Create the channel positions on the grid channel_positions_x, channel_positions_y = np.meshgrid(channel_positions_x, channel_positions_y) channel_positions_x = channel_positions_x.flatten() channel_positions_y = channel_positions_y.flatten() channel_positions = tf.stack([channel_positions_x, channel_positions_y, channel_positions_z], axis=1) def magnetic_dipole_fast_tf(R, pos, ori): """ Calculate the leadfield for a magnetic dipole in an infinite medium using TensorFlow operations. Parameters: R (Tensor): Position of the magnetic dipole (1x3 Tensor). pos (Tensor): Positions of magnetometers (Nx3 Tensor). ori (Tensor): Orientations of magnetometers (Nx3 Tensor). Returns: lf (Tensor): Leadfield matrix (Nx3 Tensor). # Example usage: R = tf.constant([1.0, 2.0, 3.0], dtype=tf.float32) pos = tf.constant([[4.0, 5.0, 6.0], [7.0, 8.0, 9.0]], dtype=tf.float32) # Example positions of magnetometers ori = tf.constant([[0.0, 0.0, 1.0], [0.0, 1.0, 0.0]], dtype=tf.float32) # Example orientations of magnetometers leadfield = magnetic_dipole_fast_tf(R, pos, ori) """ mu_0 = 4 * np.pi * 1e-7 # Permeability of free space # Modify the shape of R to be a 2D tensor with one row R = tf.reshape(R, (1, 3)) # Modify the shape of pos to be a 2D tensor with one row per magnetometer pos = tf.reshape(pos, (-1, 3)) # Calculate the distance vectors between the dipole and all magnetometers r = pos - R # Calculate the distance from the dipole to each magnetometer distance = tf.norm(r, axis=1) # Calculate the magnetic field contribution for all magnetometers dot_product = tf.reduce_sum(ori * r, axis=1) r_squared = distance**2 lf = (mu_0 / (4 * np.pi)) * (3 * dot_product[:, tf.newaxis] * r / r_squared[:, tf.newaxis]**(5/2) - ori / distance[:, tf.newaxis]**3) return lf def generate_H(channel_positions, source_positions, source_orientations): num_channels = channel_positions.shape[0] num_sources = source_positions.shape[0] # 用tf.zeros替代np.zeros初始化TensorFlow张量 H = tf.zeros((num_channels, num_sources, 3), dtype=tf.float32) for i in range(num_channels): for j in range(num_sources): channel_position = channel_positions[i] source_position = source_positions[j] source_orientation = source_orientations[j] # Calculate the leadfield for the current source and channel lf = magnetic_dipole_fast_tf(source_position, channel_position, source_orientation) # 用tf.tensor_scatter_nd_update进行索引赋值,保留梯度链 indices = tf.constant([[i, j, 0], [i, j, 1], [i, j, 2]]) H = tf.tensor_scatter_nd_update(H, indices, tf.reshape(lf, (-1,))) return H # Generating the true H matrix based on initial source positions H_true = generate_H(channel_positions, source_positions, source_orientations) # Just take the first axis - assume the source orientation is known H_true = H_true[:,:,0] num_channels = channel_positions.shape[0] C_X = np.eye(num_sources) C_n = np.eye(num_channels) * 0.01 C_Y = np.dot(np.dot(H_true, C_X), tf.transpose(H_true)) + C_n # Convert matrices to TensorFlow constants C_X_tf = tf.constant(C_X, dtype=tf.float32) C_n_tf = tf.constant(C_n, dtype=tf.float32) C_Y_tf = tf.constant(C_Y, dtype=tf.float32) # Define the optimization process learning_rate = 0.1 optimizer = tf.optimizers.Adam(learning_rate) poly_coefficients = initial_poly_coefficients # Optimization loop num_iterations = 1800 tolerance = 1e-9 for iteration in range(num_iterations): with tf.GradientTape() as tape: # Generate H matrix based on fixed x-coordinates and make an estimate for the z coordinate. The # y coordinate remains fixed, too. source_positions_z_tf = tf.math.polyval(poly_coefficients, source_positions_x) # Now we need to stack the source positions into a tensor: source_positions = tf.stack([source_positions_x, source_positions_y, source_positions_z_tf], axis=1) # Update the estimate for H, and take just the last dimension H_guess = generate_H(channel_positions, source_positions, source_orientations)[:,:,0] # Calculate the difference between the two sides of the equation difference = C_Y_tf - tf.matmul(tf.matmul(H_guess, C_X_tf), tf.transpose(H_guess)) - C_n_tf # Calculate the cost total_cost = tf.reduce_sum(tf.square(difference)) # Calculate gradients gradients = tape.gradient(total_cost, poly_coefficients) # Apply gradients optimizer.apply_gradients(zip(gradients, poly_coefficients))
内容的提问来源于stack exchange,提问作者timmr002
相关产品推荐
相关产品推荐

