You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

合并两类不同训练模型时出现结构连接异常的技术求助

Fixing Model Merging Issues: Input Tensors & Unconnected Layers

Let's break down why your merged model isn't working as expected, then walk through the corrected code step by step.

Core Problem: Manually Popping the Input Layer

Your biggest issue comes from these lines:

model_simple.layers.pop(0)
model_complexe.layers.pop(0)

Keras models rely on a intact computational graph to track input/output tensor relationships. When you pop the input layer directly, you break this graph entirely—so when you later pass input_common to the model, Keras can't properly link the new input to the remaining layers of your pre-trained models. This explains why Netron shows unconnected layers, your model size is almost identical to model_simple.h5, and converting to .pb throws a "0 tensor inputs" error.

Corrected Approach

You don't need to manually remove input layers. Keras automatically handles mapping a new input tensor to a pre-trained model as long as the shapes match. Here's the fixed code:

# Keep your existing imports for ImageDataGenerator, TensorBoard, etc.
from tensorflow.keras.models import load_model, Model
from tensorflow.keras.layers import Input, concatenate, Dense
from tensorflow.keras.optimizers import Adam

# Load and prepare pre-trained models
model_simple = load_model("model_simple.h5")
model_simple.name = 'model_simple'
for layer in model_simple.layers:
    layer.trainable = False
    layer.name = f"{layer.name}_simple"  # Ensure unique layer names to avoid graph conflicts

model_complexe = load_model("model_complexe.h5")
model_complexe.name = 'model_complexe'
for layer in model_complexe.layers:
    layer.trainable = False
    layer.name = f"{layer.name}_complexe"  # Unique names prevent silent failures

# Create shared input tensor
input_common = Input(shape=(299, 299, 3), name="input_common")

# Get outputs from both models using the shared input
model_simple_output = model_simple(input_common)
model_complexe_output = model_complexe(input_common)

# Merge outputs and build the top classification layers
x = concatenate([model_simple_output, model_complexe_output])
x = Dense(2 * NB_CLASSES, activation='relu')(x)
x = Dense(4 * NB_CLASSES, activation='relu')(x)
x = Dense(4 * NB_CLASSES, activation='relu')(x)
x = Dense(NB_CLASSES, activation='relu')(x)
# Match activation function to your task type
output = Dense(NB_CLASSES, activation='softmax')(x)  # Use softmax for multi-class; sigmoid for multi-label

# Build the final merged model
model = Model(inputs=input_common, outputs=output)

# Compile with a loss function that matches your activation
model.compile(
    optimizer=Adam(lr=0.0001, beta_1=0.9, beta_2=0.999, epsilon=1e-8, amsgrad=True),
    loss='categorical_crossentropy',  # Use binary_crossentropy if using sigmoid for multi-label
    metrics=['acc']
)

# Train and save the final model
model.fit_generator(
    train_generator,
    steps_per_epoch=NB_FIC_TRAIN // BATCH_SIZE,
    epochs=1,
    validation_data=validation_generator,
    validation_steps=NB_FIC_VAL // BATCH_SIZE,
    callbacks=[tensorboard]
)

model.save("modele_final.h5")

Key Notes for Success

  1. Layer Name Uniqueness: Your code already adds _simple/_complexe suffixes—this is critical to avoid layer name conflicts in the merged graph, which can cause silent failures that are hard to debug.
  2. Activation & Loss Match: Your original code uses sigmoid activation with categorical_crossentropy—this is a mismatch. Follow these rules:
    • Use softmax + categorical_crossentropy for multi-class tasks (mutually exclusive labels)
    • Use sigmoid + binary_crossentropy for multi-label tasks (multiple labels can apply to one sample)
  3. Input Shape Compatibility: Ensure both pre-trained models were trained on input shape (299, 299, 3)—this is already handled in your code, but double-check if you ever adjust input sizes.

After making these changes, load the saved modele_final.h5 in Netron: you should see input_common connected to both pre-trained models, their outputs feeding into the concatenate layer, and the full graph leading to the final output. Converting to .pb should also work without input tensor errors.

内容的提问来源于stack exchange,提问作者Enzo Dutra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:16:20