为何Keras Input与常规张量表现不同?附报错解析
我在Keras中构建了Embedding层和GRU层:
embedding_layer = tf.keras.layers.Embedding(5000, 256, mask_zero=True) gru_layer = tf.keras.layers.GRU(256, return_sequences=True, recurrent_initializer='glorot_uniform')
当使用numpy数组输入时:
A1 = np.random.random((64, 29)) A2 = embedding_layer(A1) A3 = gru_layer(A2) print(A1.shape, A2.shape, A3.shape)
运行正常,输出形状为(64, 29) (64, 29, 256) (64, 29, 256)。
但使用Keras Input输入时:
y2 = tf.keras.Input(shape=(64,29)) print(y2.shape) y3 = embedding_layer(y2) print(y3.shape) y4 = gru_layer(y3) print(y4.shape)
前两步输出形状为(None, 64, 29)和(None, 64, 29, 256),但后续调用GRU层时出现维度不匹配的ValueError:
ValueError: Input 0 of layer "gru_17" is incompatible with the layer: expected ndim=3, found ndim=4. Full shape received: (None, 64, 29, 256)
疑问:为何Keras Input与常规张量表现不同导致报错?且张量形状为何显示为(None, 64, 29)而非(64, 29)?
1. 维度不匹配报错的原因
Keras的Input层的shape参数定义的是单个样本的形状,而非整个批次的形状:
- 用numpy数组
(64,29)输入时,64是批次大小,29是单个样本的序列长度。Embedding层会把每个样本的每个元素映射为256维向量,输出形状为(64,29,256),符合GRU层要求的3维输入格式(batch_size, sequence_length, feature_dim)。 - 而你定义
Input(shape=(64,29))时,是将单个样本的形状设为(64,29),此时输入张量的完整形状是(None,64,29)(None代表批次大小可变)。Embedding层会对这个4维张量的最后一维做嵌入,输出变成(None,64,29,256),属于4维张量,而GRU层仅接受3维输入,因此报错。
解决方法:将Input的shape改为单个样本的序列长度,即shape=(29,),修正后代码如下:
y2 = tf.keras.Input(shape=(29,)) print(y2.shape) # 输出 (None, 29) y3 = embedding_layer(y2) print(y3.shape) # 输出 (None, 29, 256) y4 = gru_layer(y3) print(y4.shape) # 输出 (None, 29, 256)
2. 张量形状显示为(None,64,29)的原因
Input层的shape参数不包含批次维度,Keras会自动在最前面添加一个可变的批次维度(用None表示),代表模型可以接受任意大小的批次输入。而直接用numpy数组输入时,numpy数组已经包含了具体的批次维度(64),所以输出形状会直接显示该数值。
简单来说:Input层定义的是样本维度,批次维度由Keras自动补充为可变的None;numpy数组是直接传入包含批次的完整张量,因此形状显示具体的批次数值。
内容的提问来源于stack exchange,提问作者mhsnk

