You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Keras标签(Y)预处理层整合进模型?

将标签归一化整合到Keras模型中的方案

注意:你的原始特征归一化函数存在一个小问题——feature_ds被初始化为None,且dataset.map(...)的结果没有赋值给它,这会导致normalizer.adapt(feature_ds)运行报错。修正后的函数应该是:

def get_normalization_layer(name, dataset):
    # 提取指定特征的数据集
    feature_ds = dataset.map(lambda x, y: x[name])
    normalizer = layers.Normalization(axis=None)
    normalizer.adapt(feature_ds)
    return normalizer

要把标签归一化完全整合到模型中,需要同时处理训练时的标签归一化和预测时的反归一化,确保整个数据处理管道封装在模型内部。具体实现步骤如下:

1. 创建标签归一化层

和特征归一化层逻辑一致,从训练数据集提取标签并适配归一化层:

def get_label_normalization_layer(dataset):
    # 提取训练集的标签列
    label_ds = dataset.map(lambda x, y: y)
    label_normalizer = layers.Normalization(axis=None)
    label_normalizer.adapt(label_ds)
    return label_normalizer

2. 构建端到端模型(训练+预测一体化)

我们可以构建两个关联的模型:一个用于训练(接受原始特征和原始标签),另一个用于预测(仅接受原始特征,输出原始尺度的预测结果),两者共享核心网络层和归一化层:

import tensorflow as tf
from tensorflow.keras import layers

# 假设你有数值特征列表
numeric_features = ["feature1", "feature2", ...]

# 初始化特征和标签的归一化层
feature_normalizers = {feat: get_normalization_layer(feat, train_ds) for feat in numeric_features}
label_normalizer = get_label_normalization_layer(train_ds)

# ---------------------- 构建核心特征处理和预测网络 ----------------------
inputs = {}
encoded_features = []
for feat in numeric_features:
    input_layer = tf.keras.Input(shape=(1,), name=feat)
    inputs[feat] = input_layer
    # 特征归一化
    normalized_feat = feature_normalizers[feat](input_layer)
    encoded_features.append(normalized_feat)

# 拼接特征并构建预测主体网络
x = layers.concatenate(encoded_features)
x = layers.Dense(64, activation="relu")(x)
x = layers.Dense(32, activation="relu")(x)
# 输出归一化后的预测值
normalized_pred = layers.Dense(1)(x)

# ---------------------- 构建训练专用模型 ----------------------
# 训练时需要输入原始标签
label_input = layers.Input(shape=(1,), name="label")
# 模型内部对标签进行归一化
normalized_label = label_normalizer(label_input)

# 计算损失(基于归一化后的标签和预测值)
loss = tf.keras.losses.MeanSquaredError()(normalized_label, normalized_pred)

# 训练模型:输入为特征+原始标签,输出为归一化预测值
training_model = tf.keras.Model(inputs=[inputs, label_input], outputs=normalized_pred)
# 将损失添加到模型中
training_model.add_loss(loss)
# 编译模型
training_model.compile(optimizer=tf.keras.optimizers.Adam())

# ---------------------- 构建预测专用模型 ----------------------
# 将归一化的预测值反归一化,还原为原始标签尺度
denormalized_pred = label_normalizer.inverse(normalized_pred)
prediction_model = tf.keras.Model(inputs=inputs, outputs=denormalized_pred)

3. 训练与预测

  • 训练:直接传入包含原始特征和原始标签的训练数据集:
training_model.fit(train_ds.map(lambda x, y: (x, y)), epochs=10)
  • 预测:使用prediction_model直接输入原始特征,得到原始尺度的预测结果:
# 提取测试集的特征部分
test_feat_ds = test_ds.map(lambda x, y: x)
predictions = prediction_model.predict(test_feat_ds)

关键优势

  • 所有数据预处理逻辑(特征归一化、标签归一化/反归一化)都封装在模型内部,无需在外部用dataset.map()处理标签
  • 训练和预测模型共享核心层,确保预处理逻辑一致
  • 保存模型时会自动包含所有归一化层,部署时无需额外处理数据

内容的提问来源于stack exchange,提问作者araisbec

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 01:52:21