合并深度学习模型时遇KerasTensor输出计算错误问题求助
多分类任务中合并深度学习模型的输出层错误解决
问题背景与代码
构建多分类任务的两个深度学习模型并合并时,输出层出现计算错误,代码如下:
import numpy as np import tensorflow as tf from tensorflow import keras from keras import layers from tensorflow.keras.layers import Input, Dense, concatenate, Flatten from tensorflow.keras.models import Model from tensorflow.keras.utils import plot_model from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.model_selection import train_test_split # 补充缺失的导入 # Model A a_ip_img = Input(shape=(5238,1,1500), name="Input_a") al_1 = Dense(64, activation = "relu",name ="a_layer_1")(a_ip_img) al_2 = Dense(128, activation="relu",name ="a_layer_2")(al_1) al_3 = Dense(64, activation="relu",name ="a_layer_3")(al_2) al_4 = Dense(32, activation="softmax",name ="a_output_layer")(al_3) #Model B b_ip_img = Input(shape=(5238,1,1500), name="Input_b") bl_1 = Dense(64, activation="relu",name ="b_layer_1")(b_ip_img) bl_2 = Dense(32, activation = "softmax",name ="b_output_layer")(bl_1) #Merging model A and B a_b = concatenate([al_4,bl_2],name="concatenated_layer") #Final Layer output_layer = Dense(7, activation = "softmax", name = "output_layer")(a_b) #Model Definition merged = Model(inputs=[(a_ip_img,b_ip_img)],outputs=[output_layer], name = "merged_model") #Model Details merged.summary() keras.utils.plot_model(merged, "output/architecture.png", show_shapes=True) opt1 = keras.optimizers.Adam(learning_rate=0.001) merged.compile(loss='categorical_crossentropy', optimizer=opt1,metrics=['accuracy']) wordsfreq=TfidfVectorizer(max_features=1500) TF_IDF_Words1 = wordsfreq.fit_transform(your_text_data).toarray() # 假设此处有文本数据处理 Y = keras.utils.to_categorical(your_labels, num_classes=7) # 假设标签已转为独热编码 X_train, X_test, y_train, y_test = train_test_split(TF_IDF_Words1, Y, test_size=0.33, random_state=42) history=merged.fit([X_train,X_train],y_train,validation_data=(X_test,y_test),epochs=20,batch_size=128)
错误信息
AssertionError: Could not compute output KerasTensor(type_spec=TensorSpec(shape=(None, 5238, 1, 7), dtype=tf.float32, name=None), name='output_layer/Softmax:0', description="created by layer 'output_layer'")
问题分析
- 输入形状与实际数据不匹配:TF-IDF生成的文本特征是二维数组(样本数 × 1500),但代码中定义的Input形状为
(5238,1,1500),完全不符合数据维度,导致后续层的张量维度混乱。 - 高维输入未扁平化:Dense层默认仅作用于张量的最后一维,若输入是3D张量,中间层处理后仍会保留前两维,最终输出层的形状为
(None,5238,1,7),与7分类任务所需的(None,7)标签维度不匹配。 - 模型输入定义错误:
inputs=[(a_ip_img,b_ip_img)]多了一层嵌套,Keras无法正确识别输入列表。 - 中间层误用softmax:模型A和B的中间输出层使用softmax会抑制特征表达,特征融合阶段应保留relu激活的特征,仅在最终输出层使用softmax。
修正方案
以下是调整后的完整代码:
import numpy as np import tensorflow as tf from tensorflow import keras from tensorflow.keras.layers import Input, Dense, concatenate, Flatten from tensorflow.keras.models import Model from tensorflow.keras.utils import plot_model from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.model_selection import train_test_split # 假设此处有你的文本数据和标签 # your_text_data = ["文本1", "文本2", ...] # your_labels = [0, 1, 2, ...] # TF-IDF特征处理 wordsfreq = TfidfVectorizer(max_features=1500) TF_IDF_Words1 = wordsfreq.fit_transform(your_text_data).toarray() # 标签转为独热编码 Y = keras.utils.to_categorical(your_labels, num_classes=7) # Model A:适配二维TF-IDF输入 a_ip = Input(shape=(1500,), name="Input_a") al_1 = Dense(64, activation="relu", name="a_layer_1")(a_ip) al_2 = Dense(128, activation="relu", name="a_layer_2")(al_1) al_3 = Dense(64, activation="relu", name="a_layer_3")(al_2) al_4 = Dense(32, activation="relu", name="a_feature_layer")(al_3) # 中间层用relu # Model B:适配二维TF-IDF输入 b_ip = Input(shape=(1500,), name="Input_b") bl_1 = Dense(64, activation="relu", name="b_layer_1")(b_ip) bl_2 = Dense(32, activation="relu", name="b_feature_layer")(bl_1) # 中间层用relu # 合并两个模型的特征 concatenated = concatenate([al_4, bl_2], name="concatenated_layer") # 最终输出层:7分类用softmax output_layer = Dense(7, activation="softmax", name="output_layer")(concatenated) # 修正模型输入定义:移除多余嵌套 merged_model = Model(inputs=[a_ip, b_ip], outputs=output_layer, name="merged_model") # 模型编译与训练 merged_model.summary() keras.utils.plot_model(merged_model, "output/architecture.png", show_shapes=True) opt1 = keras.optimizers.Adam(learning_rate=0.001) merged_model.compile(loss='categorical_crossentropy', optimizer=opt1, metrics=['accuracy']) X_train, X_test, y_train, y_test = train_test_split(TF_IDF_Words1, Y, test_size=0.33, random_state=42) # 两个输入均为X_train(若两个模型处理相同特征) history = merged_model.fit([X_train, X_train], y_train, validation_data=([X_test, X_test], y_test), # 验证集输入也需对应两个输入 epochs=20, batch_size=128)
关键调整点
- 将Input形状改为
(1500,),匹配TF-IDF特征的实际维度。 - 移除中间层的softmax,改用relu保留特征表达能力。
- 修正模型输入列表的嵌套问题,确保Keras正确识别输入张量。
- 验证集输入调整为
[X_test, X_test],与模型的两个输入对应。
内容的提问来源于stack exchange,提问作者user1873093
相关产品推荐
相关产品推荐

