TensorFlow Keras Dot层形状不兼容错误排查求助
解决Keras自定义层与Dense层乘积的形状不兼容问题
问题复现
构建回归模型时,输出为标量Dense层与标量自定义层的乘积,输入是二维数组。自定义层仅处理输入数组的第一个元素,代码如下:
自定义层代码
class CustomLayer(tf.keras.layers.Layer): def __init__(self, *args, **kwargs): super().__init__( *args, **kwargs) self.a = tf.Variable(0.1, dtype=tf.float32, trainable=True) self.b = tf.Variable(0.1, dtype=tf.float32, trainable=True) self.rho = tf.Variable(0.1, dtype=tf.float32, trainable=True) self.m = tf.Variable(0, dtype=tf.float32, trainable=True) self.sigma = tf.Variable(0.1, dtype=tf.float32, trainable=True) def call(self, inputs): return self.a + self.b * (self.rho * (inputs[0] - self.m) + tf.math.sqrt( tf.math.square( (inputs[0] - self.m) ) + tf.math.square( self.sigma)))
函数式API代码
inputs = keras.Input(shape=(2, ), name="digits") x1 = layers.Dense(40, activation="softplus", kernel_initializer='normal')(inputs) x2 = layers.Dense(1, activation="softplus")(x1) custom = CustomLayer()(inputs) outputs = layers.dot([custom, x2], 0) # 尝试过不同axis参数均报错 model = keras.Model(inputs=inputs, outputs=outputs)
报错信息
Dot.build(self, input_shape) 134 axes = self.axes 135 if shape1[axes[0]] != shape2[axes[1]]: --> 136 raise ValueError( 137 "Incompatible input shapes: " 138 f"axis values {shape1[axes[0]]} (at axis {axes[0]}) != " 139 f"{shape2[axes[1]]} (at axis {axes[1]}). " 140 f"Full input shapes: {shape1}, {shape2}" 141 ) ValueError: Incompatible input shapes: axis values 2 (at axis 0) != None (at axis 0). Full input shapes: (2,), (None, 1)
张量形状
- x2:
<KerasTensor: shape=(None, 1) dtype=float32 (created by layer 'dense_191')>(None代表batch大小) - custom:
<KerasTensor: shape=(2,) dtype=float32 (created by layer 'custom_layer_3')>
问题原因
自定义层丢失batch维度的核心原因是索引方式错误:
Keras输入张量的形状为(batch_size, feature_dim),第一个维度是batch,第二个是特征数。代码中inputs[0]是取整个输入的第0个维度(即第一个样本的所有特征),得到的形状是(2,),完全丢失了batch维度。而我们需要的是每个样本的第一个特征,应该用inputs[:, 0]来索引,这样得到的形状是(batch_size,),保留了batch维度。
解决方案
步骤1:修正自定义层的索引逻辑
修改CustomLayer的call方法,将inputs[0]替换为inputs[:, 0],确保输出保留batch维度:
class CustomLayer(tf.keras.layers.Layer): def __init__(self, *args, **kwargs): super().__init__( *args, **kwargs) self.a = tf.Variable(0.1, dtype=tf.float32, trainable=True) self.b = tf.Variable(0.1, dtype=tf.float32, trainable=True) self.rho = tf.Variable(0.1, dtype=tf.float32, trainable=True) self.m = tf.Variable(0, dtype=tf.float32, trainable=True) self.sigma = tf.Variable(0.1, dtype=tf.float32, trainable=True) def call(self, inputs): # 取每个样本的第一个特征,形状为(batch_size,) first_feature = inputs[:, 0] return self.a + self.b * (self.rho * (first_feature - self.m) + tf.math.sqrt( tf.math.square(first_feature - self.m) + tf.math.square(self.sigma)))
步骤2:调整模型输出的乘积逻辑
修正后,custom的形状变为(None,),x2的形状是(None,1),可以通过以下两种方式完成乘积:
方式1:利用TensorFlow广播机制直接相乘
inputs = keras.Input(shape=(2, ), name="digits") x1 = layers.Dense(40, activation="softplus", kernel_initializer='normal')(inputs) x2 = layers.Dense(1, activation="softplus")(x1) custom = CustomLayer()(inputs) # 利用广播自动匹配维度,或者手动压缩x2的维度 outputs = custom * tf.squeeze(x2, axis=1) model = keras.Model(inputs=inputs, outputs=outputs)
方式2:用Dot层实现(需统一维度)
将custom扩展为(None,1)的形状,再与x2在axis=1上做点积:
inputs = keras.Input(shape=(2, ), name="digits") x1 = layers.Dense(40, activation="softplus", kernel_initializer='normal')(inputs) x2 = layers.Dense(1, activation="softplus")(x1) custom = CustomLayer()(inputs) # 将custom从(None,)转为(None,1) custom_expanded = layers.Reshape((1,))(custom) outputs = layers.dot([custom_expanded, x2], axes=1) model = keras.Model(inputs=inputs, outputs=outputs)
内容的提问来源于stack exchange,提问作者acai
相关产品推荐
相关产品推荐

