TensorFlow多标签评论分类模型训练报错求助:维度异常ValueError
在基于TensorFlow训练多标签评论分类模型时出现维度相关报错,以下是相关信息及排查方案:
模型定义代码
def mlp_model(layers, units, dropout_rate, input_shape, num_classes): model = Sequential() model.add(Embedding(max_words, 20, input_length=maxlen ,input_shape=input_shape)) model.add(Dropout(rate=dropout_rate, input_shape=input_shape)) for _ in range(layers-1): model.add(Dense(units=units, activation='relu')) model.add(Dropout(rate=dropout_rate)) model.add(GlobalMaxPool1D()) model.add(Dense(units=num_classes, activation='sigmoid')) return model
训练代码
model = mlp_model(layers = 2, units = 50,dropout_rate = 0.2, input_shape = (100,), num_classes = 9) model.compile(optimizer='adam' , loss='binary_crossentropy', metrics=[tf.keras.metrics.AUC()]) callbacks = [ ReduceLROnPlateau(), ModelCheckpoint(filepath='model-simple.h5', save_best_only=True) ] history = model.fit(X_train.astype('float32'), y_train, class_weight=class_weight, epochs=25, batch_size=32, validation_split=0.2, callbacks=callbacks)
报错信息
ValueError: Can not squeeze dim[1], expected a dimension of 1, got 0
for '{{node binary_crossentropy/weighted_loss/Squeeze}} =
SqueezeT=DT_FLOAT, squeeze_dims=[-1]' with input
shapes: [?,0].
已知数据形状
- X_train: (99909, 100)
- y_train: (99909, 9)
问题原因及解决方法
Dropout层错误指定input_shape
原代码中Embedding层后的Dropout层手动添加了input_shape=input_shape,这是完全多余的:Embedding层输出的是3D张量(batch_size, 100, 20),Dropout层会自动继承上一层的输出形状,手动指定(100,)会导致维度不匹配,后续网络层处理时出现维度坍塌,最终引发损失计算的维度错误。3D张量直接接入Dense层的逻辑问题
Embedding层输出的3D张量直接接入Dense层,会让Dense层作用于序列的每个时间步,输出形状变为(batch_size, 100, 50),虽然最后有GlobalMaxPool1D做降维,但前面的维度混乱已经埋下隐患。
修改后的模型代码
移除Dropout层的多余input_shape参数,同时调整结构先做池化再接入Dense层(更符合MLP处理文本的常规逻辑):
def mlp_model(layers, units, dropout_rate, input_shape, num_classes): model = Sequential() # 移除重复的input_shape参数,input_length已指定序列长度 model.add(Embedding(max_words, 20, input_length=maxlen)) model.add(Dropout(rate=dropout_rate)) # 先通过池化把3D张量转为2D,再接入全连接层 model.add(GlobalMaxPool1D()) for _ in range(layers-1): model.add(Dense(units=units, activation='relu')) model.add(Dropout(rate=dropout_rate)) model.add(Dense(units=num_classes, activation='sigmoid')) return model
额外检查:确保函数外部的max_words和maxlen变量已正确定义,值与输入数据的词汇量、序列长度匹配。
内容的提问来源于stack exchange,提问作者AlirezaTomari

