You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Keras的二进制向量转0-100标量预测模型优化咨询

问题描述

我需要用神经网络从250维二进制输入向量预测0到100之间的标量值,现有1000组输入输出数据:

>>>in_.shape
(1000, 250)

>>>in_[0]
array([1, 0, 1, 1, 1, 1, 1, 1, ...])

>>>out.shape
(1000,)

>>>out[0]
64.46677867594474

编写的Keras模型如下,但未正常工作:

import numpy as np
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers


in_ = np.load('input.npy')
out = np.load('output.npy')

model = keras.Sequential([
        keras.Input(shape=(250,)),
        layers.Dense(1000, activation='relu'),
        layers.Dense(1000, activation='relu'),
        layers.Dense(250, activation='relu'),
        layers.Dense(1, activation='linear')])

model.compile(loss='binary_crossentropy', optimizer='adam',
              metrics=['accuracy'])

model.fit(in_, out, batch_size=100, epochs=5, validation_split=0.1)

训练日志:

Epoch 1/5
9/9 [==============================] - 2s 54ms/step - loss: -240.3843 - accuracy: 0.0000e+00 - val_loss: -54.9291 - val_accuracy: 0.0000e+00
Epoch 2/5
9/9 [==============================] - 0s 18ms/step - loss: -311.0255 - accuracy: 0.0000e+00 - val_loss: -54.9291 - val_accuracy: 0.0000e+00
Epoch 3/5
9/9 [==============================] - 0s 20ms/step - loss: -311.0255 - accuracy: 0.0000e+00 - val_loss: -54.9291 - val_accuracy: 0.0000e+00
Epoch 4/5
9/9 [==============================] - 0s 18ms/step - loss: -311.0254 - accuracy: 0.0000e+00 - val_loss: -54.9291 - val_accuracy: 0.0000e+00
Epoch 5/5
9/9 [==============================] - 0s 18ms/step - loss: -311.0255 - accuracy: 0.0000e+00 - val_loss: -54.9291 - val_accuracy: 0.0000e+00
改进方案

1. 修正损失函数与评估指标

当前任务是回归任务(预测连续标量),但误用了分类任务的binary_crossentropy损失和accuracy指标,这是核心错误:

  • 损失函数替换为回归专用的mean_squared_error(均方误差)或mean_absolute_error(平均绝对误差)
  • 评估指标改用mean_absolute_error,accuracy对回归任务无意义

修改后的编译代码:

model.compile(loss='mean_squared_error', optimizer='adam',
              metrics=['mean_absolute_error'])

2. 缩小模型规模,抑制过拟合

输入仅250维,却使用两个1000维全连接层,对于1000样本的数据集来说模型过于庞大,极易过拟合:

  • 减少隐藏层神经元数量,比如调整为256→128→64的层级结构
  • 添加Dropout层随机失活部分神经元,降低过拟合风险

调整后的模型结构:

model = keras.Sequential([
        keras.Input(shape=(250,)),
        layers.Dense(256, activation='relu'),
        layers.Dropout(0.2),
        layers.Dense(128, activation='relu'),
        layers.Dropout(0.2),
        layers.Dense(64, activation='relu'),
        layers.Dense(1, activation='linear')])

3. 调整训练策略,确保模型收敛

原训练仅5轮,模型未充分学习:

  • 增加训练轮数至50-100轮
  • 添加EarlyStopping回调,当验证损失连续多轮无下降时自动停止训练,保留最优权重

示例代码:

from tensorflow.keras.callbacks import EarlyStopping

early_stop = EarlyStopping(monitor='val_loss', patience=5, restore_best_weights=True)
model.fit(in_, out, batch_size=100, epochs=100, validation_split=0.1, callbacks=[early_stop])

4. 可选:输出归一化加速收敛

将输出值缩放到0-1区间(如除以100),可帮助模型更快收敛,预测时再还原:

# 归一化输出
out_scaled = out / 100.0
model.fit(in_, out_scaled, ...)

# 预测时还原结果
predictions = model.predict(test_in) * 100.0

内容的提问来源于stack exchange,提问作者oakca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 15:05:16