You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Keras构建消费得分预测模型:损失下降但准确率低且过拟合求助

问题:基于200条数据预测Spending Score的模型异常问题

我用200条数据构建预测**Spending Score (1-100)**的机器学习模型,执行了以下操作:

  • 对Genre列做独热编码
  • 循环使用MinMaxScaler单独缩放各特征列
  • 用Keras搭建神经网络

遇到的问题:

  • 损失函数快速下降,但准确率极低
  • 增加网络层数后,训练准确率提升,但测试集表现极差(严重过拟合)

附上代码:

df=df.drop('CustomerID',axis=1)
df.head()
ohe=preprocessing.OneHotEncoder()
feature_array=ohe.fit_transform(df[['Genre']]).toarray()
feature_labels=np.array(ohe.categories_).ravel()

features=pd.DataFrame(feature_array,columns=feature_labels)
df=pd.concat([df,features],axis=1)
df=df.drop('Genre',axis=1) 

df=df[['Male','Female','Age','Annual Income (k$)','Spending Score (1-100)']]
df.head()

scaler = preprocessing.MinMaxScaler()

for col in df.columns:
    if col != 'Spending Score (1-100)':
        df[col] = scaler.fit_transform(df[[col]])

df.head()

from sklearn.model_selection import train_test_split
x_train,x_test,y_train,y_test=train_test_split(df.drop('Spending Score (1-100)',axis=1),df['Spending Score (1-100)'],test_size=0.2)

model=keras.Sequential([
  keras.layers.Dense(50,input_shape=(4,),activation='relu'),
  keras.layers.Dense(75,activation='relu'),
  keras.layers.Dense(100,activation='sigmoid')

])

model.compile(optimizer='adam',loss='sparse_categorical_crossentropy',metrics=['accuracy'])
                                    
model.fit(x_train,y_train,epochs=1000,batch_size=10)

问题分析与解决步骤

1. 核心错误:任务类型判断混淆

Spending Score是1-100的连续数值,属于回归任务,但你用了分类任务的sparse_categorical_crossentropy损失函数:

  • 分类损失会把每个分数当成独立类别,200条数据要分100类,样本量严重不足,直接导致准确率极低
  • 修正:损失函数换成mse(均方误差),最后一层激活函数用linear(回归不需要非线性激活)

2. 独热编码冗余问题

Genre列独热编码后生成Male和Female两列,存在完全共线性(Male=1-Female),增加模型冗余,建议只保留其中一列(比如Male)

3. 数据缩放的错误操作

你用同一个MinMaxScaler对每列单独fit_transform,会导致数据泄露:

  • 正确做法:仅在训练集上fit scaler,再用同一个scaler对训练集和测试集做transform
  • 无需循环单独缩放,直接对所有特征列一次性处理

4. 过拟合的解决措施

200条数据属于小样本,你的网络神经元数量过多,容易过拟合,可做以下调整:

  • 大幅减少神经元数量(比如第一层20,第二层10)
  • 添加Dropout层抑制过拟合
  • 减少训练epochs,配合早停机制(EarlyStopping),避免模型在训练集过度拟合

修改后的完整代码

import pandas as pd
import numpy as np
from sklearn import preprocessing
from sklearn.model_selection import train_test_split
from keras.models import Sequential
from keras.layers import Dense, Dropout
from keras.callbacks import EarlyStopping

# 数据预处理:简化Genre编码
df = df.drop('CustomerID', axis=1)
df['Male'] = (df['Genre'] == 'Male').astype(int)
df = df.drop('Genre', axis=1)

# 特征与标签分离
X = df.drop('Spending Score (1-100)', axis=1)
y = df['Spending Score (1-100)']

# 划分训练测试集
x_train, x_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# 数据缩放:仅在训练集fit,避免数据泄露
scaler = preprocessing.MinMaxScaler()
x_train = scaler.fit_transform(x_train)
x_test = scaler.transform(x_test)

# 搭建回归模型
model = Sequential([
    Dense(20, input_shape=(3,), activation='relu'),
    Dropout(0.2),
    Dense(10, activation='relu'),
    Dropout(0.2),
    Dense(1, activation='linear')  # 回归任务用linear激活
])

# 编译模型:回归用mse损失
model.compile(optimizer='adam', loss='mse', metrics=['mae'])

# 早停回调:监控验证集损失,自动停止训练并保留最优权重
early_stop = EarlyStopping(monitor='val_loss', patience=10, restore_best_weights=True)

# 训练模型
model.fit(x_train, y_train, 
          epochs=200, 
          batch_size=10, 
          validation_data=(x_test, y_test),
          callbacks=[early_stop])

额外提示

  • 回归任务无需关注准确率,重点看MAE(平均绝对误差)或MSE(均方误差),数值越小模型效果越好
  • 小样本场景下,可先尝试线性回归、随机森林回归等简单模型做基线,再对比神经网络效果,简单模型往往更稳定

内容的提问来源于stack exchange,提问作者cheenuz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 01:02:50