You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras二分类模型始终预测同一类的问题排查与解决问询

Keras模型全预测单一类别问题排查与解决

问题现象

训练完成后模型完全预测负类,混淆矩阵如下:

PositiveNegative
Positive0 (TP)21719 (FN)
Negative0 (FP)22620 (TN)

代码模块

数据准备

data = purchase_data.copy()
labelencoder = LabelEncoder()
target_sum = 120
data.loc[data['sales'] <= target_sum, 'sales'] = False
data.loc[data['sales'] > target_sum, 'sales'] = True

print("\nColumn Names & formatting:\n")
for col in data.columns.values.tolist():
    if data[col].dtype == "object" or data[col].dtype == "bool":
        print("{:<30}".format(col), ":", "{:<30}".format(str(data[col].dtype)) , "Formatting to LabelEncoding")
        data[col] = labelencoder.fit_transform(data[col])
    else:
        print("{:<30}".format(col), ":", "{:<30}".format(str(data[col].dtype)) , "No formatting required.")

# Converting datetime to float
data['accessed_date'] = data['accessed_date'].apply(lambda x: x.timestamp())

array = data.values 
class_column = 'sales' # The column I want to predict

X = np.delete(array, data.columns.get_loc(class_column), axis=1) # Removing class_column column
Y = array[:,data.columns.get_loc(class_column)] # Selecting class_column column
Y = Y[:, np.newaxis] # Resetting the shape value

# Normalizing the input values (excluding the class value)
scaler = preprocessing.Normalizer().fit(X)
X = scaler.transform(X)

数据集拆分

seed = 1
X_train, X_test, Y_train, Y_test  = train_test_split(X, Y, test_size=0.33, random_state=seed, shuffle = True, stratify=(Y))

神经网络构建

tf.random.set_seed(seed)

# Building the neural network
modeldl = Sequential()

modeldl.add(Dense(64, input_dim=X.shape[1], activation='relu', kernel_initializer=he_normal()))
modeldl.add(Dropout(0.2))

modeldl.add(Dense(32, activation='relu', kernel_initializer=he_normal()))
modeldl.add(Dropout(0.2))

modeldl.add(Dense(1, activation='sigmoid', kernel_initializer=he_normal()))

# Compile model
optimizer = tf.keras.optimizers.Adam(learning_rate=1e-04)
modeldl.compile(loss='binary_crossentropy', optimizer=optimizer, metrics=['acc'])

results = modeldl.fit(X_train, Y_train, epochs=80, batch_size=1000, verbose=1)

已尝试方法

  • 调整超参数
  • 预测其他列而非'sales'
  • 调整网络规模
  • 固定学习率
  • 更换激活函数
  • 添加/移除Dropout层

补充信息

数据集两类样本占比约各50%,使用电商网站日志数据集,已移除与退货相关的列和行。

解决办法

  1. 核对标签编码逻辑
    用LabelEncoder处理布尔型sales列后,确认原正类(sales>120)对应的编码值是1还是0。如果正类被编码为0,模型默认0.5的阈值会导致预测逻辑反转;同时打印模型对训练集的预测概率,看是否所有值都远低于0.5。

  2. 替换归一化方式
    当前使用的Normalizer是对单样本做归一化,对多数分类任务适配性差。换成标准化或最小最大缩放,且严格只在训练集拟合避免数据泄露:

    # 用StandardScaler示例
    scaler = preprocessing.StandardScaler().fit(X_train)
    X_train = scaler.transform(X_train)
    X_test = scaler.transform(X_test)
    
  3. 调整训练策略

    • 调整学习率:将1e-4提升至5e-4或1e-3,同时将训练轮数增加到150-200,观察损失曲线是否下降。
    • 减小批次大小:把batch_size从1000改为128或256,让模型权重更新更频繁。
    • 加入验证监控:训练时添加validation_data=(X_test, Y_test),确认验证集指标是否随训练迭代提升。
  4. 排查特征有效性

    • 计算特征与标签的相关性:用data.corr()查看sales与其他特征的相关系数,如果所有特征相关性极低,说明当前特征无法区分两类,需要重新构造或选择特征。
    • 删除冗余特征:移除常数、近似常数或完全无区分度的特征,减少模型噪声。
  5. 调整模型细节

    • 更换输出层初始化:将最后一层的he_normal()换成glorot_uniform(),更适合sigmoid输出层;或直接使用默认初始化。
    • 尝试类别权重:即使数据集均衡,也可加入class_weight='balanced'参数,帮助模型关注易忽略的类别。
  6. 验证数据处理正确性

    • 检查维度匹配:打印X.shape和Y.shape,确保输入特征数与模型input_dim一致,标签维度符合要求。
    • 确认训练集分布:用np.bincount(Y_train.flatten())检查训练集两类样本占比是否接近50%,避免分层抽样失效。

内容的提问来源于stack exchange,提问作者Syrus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 06:33:08