You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras中Concatenate层报错:添加的层必须是Layer实例

问题解决:TypeError 拼接层使用错误

报错信息

TypeError: The added layer must be an instance of class Layer. Received: layer=KerasTensor(type_spec=TensorSpec(shape=(None, 80, 300), dtype=tf.float32, name=None), name='concatenate_15/concat:0', description="created by layer 'concatenate_15'") of type <class 'keras.engine.keras_tensor.KerasTensor'>

问题原因

你用Sequential模型直接添加Concatenate(axis=1)([image_model.output, caption_model.output])的结果,这是错误的。Sequential是线性堆叠层的结构,add方法只能接收Layer类实例,而你这样调用Concatenate返回的是一个KerasTensor(张量),不是层,所以触发类型错误。

另外你尝试的Concatenate([image_model, caption_model])也不对,Concatenate层需要接收的是张量输入,不是模型实例。

解决方案

改用Keras函数式API构建模型,因为你的场景是多输入(图像特征+文本序列)的分支结构,Sequential不适合处理这种情况。修改后的完整代码如下:

embedding_size = 300

# 原模型可保留(若无用可删除,此处保留你原有子模型)
model = Sequential()
model.add(
    Embedding(
        input_dim=embedding_size,
        output_dim=100,
        weights=None,
        trainable=True))
model.summary()

# 图像特征处理子模型
image_model = Sequential([
        Dense(embedding_size, input_shape=(2048,), activation='relu'),
        RepeatVector(max_len)
    ])

# 文本序列处理子模型
caption_model = Sequential([
        Embedding(vocab_size, embedding_size, input_length=max_len),
        LSTM(256, return_sequences=True),
        TimeDistributed(Dense(300))
    ])

image_model.summary()
caption_model.summary()

# -------------------------- 关键修改部分 --------------------------
# 使用函数式API构建组合模型
# 定义两个输入节点:图像特征输入、文本序列输入
image_input = tf.keras.Input(shape=(2048,))
caption_input = tf.keras.Input(shape=(max_len,))

# 获取两个子模型的输出张量
image_embedding = image_model(image_input)
caption_embedding = caption_model(caption_input)

# 拼接两个张量(正确使用Concatenate的方式)
concat_tensor = tf.keras.layers.Concatenate(axis=1)([image_embedding, caption_embedding])

# 后续层处理
x = tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(256, return_sequences=False))(concat_tensor)
x = tf.keras.layers.Dense(vocab_size)(x)
output = tf.keras.layers.Activation('softmax')(x)

# 创建最终模型,明确指定输入和输出
tmodel = tf.keras.Model(inputs=[image_input, caption_input], outputs=output)
tmodel.summary()

说明

函数式API支持多输入、多输出以及任意复杂的网络分支,非常适合图像字幕生成这类需要融合不同模态特征的任务。构建时需要先定义输入节点,再将张量依次传入各个层,最后通过Model类封装输入和输出得到完整模型。

内容的提问来源于stack exchange,提问作者Venkat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 06:25:21