You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让TensorFlow网络在大型恒星分类数据集上发生过拟合?

无法让神经网络过拟合,求解决方法

背景

我正在遵循François Chollet所著《Deep Learning with Python》中的网络设计步骤完成学位期末项目,当前阶段需要构建足够大的网络使其在数据集上发生过拟合,以此确定网络的最大容量,作为后续优化时的网络规模上限,但始终无法实现过拟合。

任务与数据集

使用恒星分类数据集完成两类分类任务:

  • 光谱类型:共100类
  • 光度类型:共10类
    例如太阳的分类为'G2V',对应光谱类型'G2'、光度类型'V'。输入为3维的观测特征数据。

模型与训练情况

我构建了双分支并行全连接网络,代码如下:

# Create our input layer:
input = keras.Input(shape=(3), name='observation_data')

# Build our spectral class
s_class_branch = layers.Dense(100000, activation='relu', name = 's_class_branch_dense_1')(input)
s_class_branch = layers.Dense(500, activation='relu', name = 's_class_branch_dense_2')(s_class_branch)

# Spectral class prediction
s_class_prediction = layers.Dense(100, 
                                  activation='softmax', 
                                  name='s_class_prediction')(s_class_branch)

# Build our luminosity class
l_class_branch = layers.Dense(100000, activation='relu', name = 'l_class_branch_dense_1')(input)
l_class_branch = layers.Dense(500, activation='relu', name = 'l_class_branch_dense_2')(l_class_branch)

# Luminosity class prediction
l_class_prediction = layers.Dense(10, 
                                  activation='softmax', 
                                  name='l_class_prediction')(l_class_branch)

# Now we instantiate our model using the layer setup above
scaled_model = Model(input, [s_class_prediction, l_class_prediction])

optimizer = keras.optimizers.RMSprop(learning_rate=0.004)

scaled_model.compile(optimizer=optimizer,
              loss={'s_class_prediction':'categorical_crossentropy',
                    'l_class_prediction':'categorical_crossentropy'},
              metrics=['accuracy'])

logdir = os.path.join("logs", "2raw100k")
tensorboard_callback = tf.keras.callbacks.TensorBoard(logdir, histogram_freq=1)

scaled_model.fit(
    input_data,{
        's_class_prediction':spectral_targets,
        'l_class_prediction':luminosity_targets
        },
    epochs=20, 
    batch_size=1000,
    validation_split=0.0,
    callbacks=[tensorboard_callback])

训练配置细节:

  • 目标标签已做独热编码,每个类别对应一个输出节点
  • 尝试过1到4层隐藏层,维度从500到100000(直至耗尽RAM)
  • 试过添加归一化层,但无效果
  • 训练轮次曾增至100轮,损失曲线显示训练损失持续缓慢下降,但始终未出现过拟合迹象

求助

我该如何调整才能让网络更好地拟合数据,出现过拟合?

内容的提问来源于stack exchange,提问作者Biscuit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 01:31:20