TensorFlow 2使用gather或boolean_mask后张量第一维度变为None问题
TensorFlow 2自定义损失中
gather索引维度异常问题解决方案 问题表现
仅在Eager模式(如Keras自定义损失函数内部)使用gather、boolean_mask等索引操作时,会出现以下维度不一致问题:
- 以张量作为索引向量时,操作后返回张量的第一维度静态形状显示为
None - 以Python原生列表作为索引时,返回张量第一维度静态形状正常,等于
len(indices)
原因说明
这是TensorFlow静态图形状推断的固有特性,并非bug:
Keras编译模型时会先构建静态计算图,Python列表是编译期常量,形状可以直接被推断;而张量索引的长度是运行时动态值,静态图阶段无法确定,所以静态形状属性.shape会显示第一维为None,但张量实际运行时的动态形状是完全正确的,不会影响后续计算逻辑。
修复方案
方案1:使用动态形状获取实际维度
不要依赖张量的静态形状属性.shape,改用tf.shape()接口获取运行时的实际维度:
# 修改损失函数内的打印逻辑 tf.print(tf.shape(tf.gather(y_true, [0, 1], axis=0))[0]) # 输出2 tf.print(tf.shape(tf.gather(y_true, ff, axis=0))[0]) # 实际运行也会输出2
方案2:手动绑定静态形状
如果后续逻辑依赖静态形状推断(比如某些层的形状校验),可以用tf.ensure_shape手动绑定你明确已知的形状:
ff = tf.where([True, True, False , False])[:, 0] gathered_y = tf.gather(y_true, ff, axis=0) # 已知ff长度为2,手动绑定第一维形状 gathered_y = tf.ensure_shape(gathered_y, (2, y_true.shape[1], y_true.shape[2]))
方案3:强制Eager执行
如果不需要静态图优化加速,可以在model.compile时指定run_eagerly=True,全程走Eager模式,静态形状会和运行时动态形状完全一致:
model.compile(loss=cutsom_gan_loss_env(model), optimizer=optimizer1, metrics=None, run_eagerly=True)
复现代码(原问题可复现版本)
import tensorflow as tf from tensorflow import keras from tensorflow.keras import Model from tensorflow.keras.layers import Input, Dense, Reshape from tensorflow.keras.datasets import mnist def cutsom_gan_loss_env(model): def custom_loss(y_true,y_pred): ff = tf.where([True, True, False , False])[:, 0] with tf.GradientTape(persistent=True) as tape: tf.print(tf.gather(y_true, [0, 1], axis=0).shape) tf.print(tf.gather(y_true, ff, axis=0).shape) tape.watch(y_true) yy = model(y_true) d_yy = tape.gradient(yy,y_true) des_loss = tf.reduce_mean(d_yy) return des_loss return custom_loss def main_(): n_hidden_units = 5 num_lay = 3 kernel_init = keras.initializers.RandomUniform(-0.1, 0.1) (x_train, y_train), _ = mnist.load_data() x_train = tf.cast(x_train,tf.float32)/255. inputs = Input(x_train.shape[1:]) x = Dense(n_hidden_units,kernel_initializer=kernel_init, activation='sigmoid' )(inputs) for _ in range(num_lay): x = Dense(n_hidden_units,kernel_initializer=kernel_init, activation='sigmoid', )(x) outputs =Reshape(x_train.shape[1:])(Dense(x_train.shape[1], kernel_initializer=kernel_init, activation='softmax')(x)) model = Model(inputs=inputs, outputs=outputs) model.summary() optimizer1 = keras.optimizers.Adam(beta_1=0.9, beta_2=0.999, epsilon=None, decay=0.0, amsgrad=True) model.compile(loss=cutsom_gan_loss_env(model), optimizer=optimizer1, metrics=None) model.fit(x_train, x_train , batch_size=1000, epochs=1, shuffle=False) if __name__=='__main__': main_()
内容的提问来源于stack exchange,提问作者Benny K
相关产品推荐
相关产品推荐

