You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

类别为数字的四分类任务中调整Keras模型输出层并实现One-Hot编码的技术咨询

解决方案:修改最后一层并实现One-Hot编码(不改动类别名称)

Absolutely! You can make this change without modifying your original class names (0, 200, 250, 300) and implement one-hot encoding correctly. Here's a step-by-step breakdown:

1. 修改模型最后一层

Your idea to change the final Dense layer to Dense(4, activation='softmax') is perfect—since you only have 4 distinct classes, there's no need for 301 neurons (that was a mistake assuming class values map directly to neuron indices). Update your model like this:

model = tf.keras.models.Sequential()
model.add(tf.keras.layers.Flatten(input_shape=(n_size,n_features)))
model.add(tf.keras.layers.Dense(64, activation='relu'))
model.add(tf.keras.layers.Dense(32, activation='relu'))
# 修改为4个神经元,对应4个类别
model.add(tf.keras.layers.Dense(4, activation='softmax'))

2. 对标签进行One-Hot编码(不改动原始类别名称)

Since your original labels are 0, 200, 250, 300, you first need to map these values to 0-3 integer indices (required for one-hot encoding), then convert them to one-hot vectors. You don't need to rename the classes—this is just a preprocessing step:

import numpy as np

# 创建类别到索引的映射(保留原始类别名称)
class_map = {0: 0, 200: 1, 250: 2, 300: 3}

# 转换训练和测试标签为索引
y_train_indices = np.array([class_map[label] for label in y_train])
y_test_indices = np.array([class_map[label] for label in y_test])

# 转换为One-Hot编码
y_train_onehot = tf.one_hot(y_train_indices, depth=4)
y_test_onehot = tf.one_hot(y_test_indices, depth=4)

3. 调整模型编译配置

Since you're now using one-hot encoded labels, you need to switch the loss function from SparseCategoricalCrossentropy to CategoricalCrossentropy:

model.compile(
    optimizer=tf.keras.optimizers.SGD(),
    loss=tf.keras.losses.CategoricalCrossentropy(),
    metrics=['accuracy']
)

4. 训练模型

When training, use the one-hot encoded labels instead of the original sparse labels:

model.fit(X_train, y_train_onehot, epochs=10, batch_size=32, validation_data=(X_test, y_test_onehot))

为什么这可行?

Your original class names (0,200,250,300) are just semantic identifiers—they don't need to correspond to the neuron indices in the final layer. By mapping them to 0-3 indices, you're telling the model to treat each unique class as a separate category, without altering the original class values themselves.

内容的提问来源于stack exchange,提问作者Lwz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 16:13:13