You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Keras中将三个训练好的二分类模型合并为单多分类模型

合并预训练二分类模型为多分类模型

问题说明

我有三个已训练完成的二分类模型,输出层均采用sigmoid激活函数:

  • 第一个模型输出0到1的概率值,判断图像是否为数字ZERO;
  • 第二个模型输出0到1的概率值,判断图像是否为数字ONE;
  • 第三个模型输出0到1的概率值,判断图像是否为数字TWO。

由于重新训练多分类模型耗时过长,希望复用现有二分类模型的隐藏层特征,将三个模型合并为一个输入形状为(28,28)、输出形状为(3)的多分类模型,且尽可能无需重新训练。

现有代码

模型加载代码

model_0 = init_binary_classification_model((28,28))
model_0.load_weights('trained_weight_of_binary_classification_to_check_whether_image_is_zero.h5')

model_1 = init_binary_classification_model((28,28))
model_1.load_weights('trained_weight_of_binary_classification_to_check_whether_image_is_one.h5')

model_2 = init_binary_classification_model((28,28))
model_2.load_weights('trained_weight_of_binary_classification_to_check_whether_image_is_two.h5')

模型初始化函数

def init_binary_classification_model(input_shape=(28,28)):
  input_layer = Input(shape=input_shape)
  tensor = Flatten()(input_layer)
  tensor = Dense(16, activation='relu')(tensor)
  tensor = Dense(8, activation='relu')(tensor)
  output_layer = Dense(1, activation='sigmoid')(tensor)

  return Model(inputs=input_layer, outputs=output_layer)

解决方案

我们可以通过复用现有模型的输出,将三个独立的sigmoid概率值拼接后做归一化,得到符合多分类要求的概率分布(输出和为1),全程无需重新训练权重:

from tensorflow.keras.layers import Input, Concatenate, Lambda
from tensorflow.keras.models import Model
import tensorflow.keras.backend as K

# 创建共享输入层
input_layer = Input(shape=(28,28))

# 获取三个二分类模型的输出
output_0 = model_0(input_layer)
output_1 = model_1(input_layer)
output_2 = model_2(input_layer)

# 拼接三个单维度输出为3维张量
concatenated_outputs = Concatenate(axis=-1)([output_0, output_1, output_2])

# 定义归一化函数,将三个概率值转换为和为1的多分类概率分布
def normalize_probabilities(x):
    return x / K.sum(x, axis=-1, keepdims=True)

# 应用归一化层
multi_class_output = Lambda(normalize_probabilities)(concatenated_outputs)

# 构建最终多分类模型
multi_class_model = Model(inputs=input_layer, outputs=multi_class_output)

# 验证模型结构
multi_class_model.summary()

说明

  • 该模型完全复用原有三个二分类模型的所有权重,无需重新训练;
  • 输出层的3个值分别对应图像为ZERO、ONE、TWO的概率,且三个值的和为1,符合多分类任务的输出要求;
  • 若需要更精准的概率校准,也可以在拼接后添加一个无激活的Dense(3)层并进行少量微调,但这一步不是必须的。

内容的提问来源于stack exchange,提问作者Muhammad Ikhwan Perwira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 13:49:59