You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在TensorFlow数据流水线内完成目标标签编码并适配模型训练?

如何在TensorFlow数据流水线内完成目标标签编码并适配模型训练?

我知道这种处理方式对当前数据来说并不高效、也不是最优解,纯粹是为了学习、搞懂TensorFlow的相关操作才这么折腾的。

我手上有个包含4个数值特征,以及3种不同字符串类别标签的DataFrame。我想搭建一套TensorFlow数据流水线:接收包含原始特征和字符串标签的tf.Dataset,同时处理特征和标签,再把处理好的数据喂给函数式模型。

我知道用Sklearn很容易在流水线外搞定这些,但就是想搞清楚——能不能完全用TensorFlow工具在流水线内实现这整套流程?


我试过的最佳方案

1. 创建数据集

feats_df = iris_df.drop(labels='Species', axis=1)
lbl_df = iris_df['Species']
# 这里先跳过划分训练集、测试集的步骤
train_ds = tf.data.Dataset.from_tensor_slices((feats_df, lbl_df))

2. 创建预处理层

# 归一化特征
normalizer = tf.keras.layers.Normalization()
features_ds = train_ds.map(lambda x, y: x)  # 只提取数据集中的特征部分
normalizer.adapt(features_ds)

# 对标签做独热编码
oh_encoder = tf.keras.layers.StringLookup(output_mode="one_hot")
labels_ds = train_ds.map(lambda x, y: y)  # 只提取数据集中的标签部分
oh_encoder.adapt(labels_ds)

3. 定义函数式模型

raw_features = tf.keras.Input(shape=(4,), name="Feature Input")
raw_labels = tf.keras.Input(shape=(1,), name="Label Input")

# 分别处理特征和标签
normalized_features = normalizer(raw_features)
encoded_labels = oh_encoder(raw_labels)
# 拼接处理后的特征和标签
preprocessed_inputs = tf.keras.layers.concatenate([normalized_features, encoded_labels], axis=1)

# 搭建神经网络主体
x = tf.keras.layers.Dense(units=16, activation="relu", name="Hidden1")(preprocessed_inputs)
x = tf.keras.layers.Dense(units=8, activation="relu", name="Hidden2")(x)
output = tf.keras.layers.Dense(units=4, activation="softmax", name="Output")(x)

# 定义完整模型
model1 = tf.keras.Model(inputs=[raw_features, raw_labels], outputs=output, name="Model1")
model1.compile(
    optimizer='adam',
    loss={"Output": tf.keras.losses.CategoricalCrossentropy()},
    metrics={"Output":[tf.keras.metrics.Accuracy()]}
)

我想实现的逻辑很清晰:把未处理的特征和标签分别传给对应的预处理层,处理完之后拼接起来,再送入神经网络训练。


遇到的问题

可惜这个方案跑不通。我在Stack Overflow上找了一堆相关帖子,但都没得到完整的解答。折腾了好几个小时后,我终于搞懂了怎么给model.fit()传数据才能避免基础报错:

# 这里先省略验证数据等部分
model1.fit(
    x=train_ds.map(lambda x,y: ({"Feature Input":x, "Label Input":y}, y)),
    epochs=2)

但新的问题又来了:损失函数拿到的是元组里的第二个元素y——也就是未处理的字符串标签,根本没法计算损失值。

尝试自定义损失函数

我还试过写一个自定义损失函数,想先把字符串标签编码后再传给TensorFlow的原生损失函数:

def custom_loss(y_true, y_pred):
    # 把原始字符串标签编码为可计算的格式
    y_true_encoded = oh_encoder(y_true)
    # 计算分类交叉熵损失
    loss = tf.keras.losses.categorical_crossentropy(y_true_encoded, y_pred)
    return loss

但结果还是报错,错误信息是Cast string to float is not supported,看起来标签还是没被正确编码,这部分对我来说有点超纲,实在搞不懂问题出在哪。


问题重申

所以我想请教:有没有办法把StringLookup层输出的独热编码标签直接传给损失函数?或者有没有完全不同的实现思路能达成我的需求?

如果这条路走不通,那只用TensorFlow工具(不用Sklearn)的话,怎么编码标签才能更适配超大规模数据集呢?

备注:内容来源于stack exchange,提问作者Bob1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 09:49:33