Keras二分类模型始终预测同一类的问题排查与解决问询
Keras模型全预测单一类别问题排查与解决
问题现象
训练完成后模型完全预测负类,混淆矩阵如下:
| Positive | Negative | |
|---|---|---|
| Positive | 0 (TP) | 21719 (FN) |
| Negative | 0 (FP) | 22620 (TN) |
代码模块
数据准备
data = purchase_data.copy() labelencoder = LabelEncoder() target_sum = 120 data.loc[data['sales'] <= target_sum, 'sales'] = False data.loc[data['sales'] > target_sum, 'sales'] = True print("\nColumn Names & formatting:\n") for col in data.columns.values.tolist(): if data[col].dtype == "object" or data[col].dtype == "bool": print("{:<30}".format(col), ":", "{:<30}".format(str(data[col].dtype)) , "Formatting to LabelEncoding") data[col] = labelencoder.fit_transform(data[col]) else: print("{:<30}".format(col), ":", "{:<30}".format(str(data[col].dtype)) , "No formatting required.") # Converting datetime to float data['accessed_date'] = data['accessed_date'].apply(lambda x: x.timestamp()) array = data.values class_column = 'sales' # The column I want to predict X = np.delete(array, data.columns.get_loc(class_column), axis=1) # Removing class_column column Y = array[:,data.columns.get_loc(class_column)] # Selecting class_column column Y = Y[:, np.newaxis] # Resetting the shape value # Normalizing the input values (excluding the class value) scaler = preprocessing.Normalizer().fit(X) X = scaler.transform(X)
数据集拆分
seed = 1 X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.33, random_state=seed, shuffle = True, stratify=(Y))
神经网络构建
tf.random.set_seed(seed) # Building the neural network modeldl = Sequential() modeldl.add(Dense(64, input_dim=X.shape[1], activation='relu', kernel_initializer=he_normal())) modeldl.add(Dropout(0.2)) modeldl.add(Dense(32, activation='relu', kernel_initializer=he_normal())) modeldl.add(Dropout(0.2)) modeldl.add(Dense(1, activation='sigmoid', kernel_initializer=he_normal())) # Compile model optimizer = tf.keras.optimizers.Adam(learning_rate=1e-04) modeldl.compile(loss='binary_crossentropy', optimizer=optimizer, metrics=['acc']) results = modeldl.fit(X_train, Y_train, epochs=80, batch_size=1000, verbose=1)
已尝试方法
- 调整超参数
- 预测其他列而非'sales'
- 调整网络规模
- 固定学习率
- 更换激活函数
- 添加/移除Dropout层
补充信息
数据集两类样本占比约各50%,使用电商网站日志数据集,已移除与退货相关的列和行。
解决办法
核对标签编码逻辑
用LabelEncoder处理布尔型sales列后,确认原正类(sales>120)对应的编码值是1还是0。如果正类被编码为0,模型默认0.5的阈值会导致预测逻辑反转;同时打印模型对训练集的预测概率,看是否所有值都远低于0.5。替换归一化方式
当前使用的Normalizer是对单样本做归一化,对多数分类任务适配性差。换成标准化或最小最大缩放,且严格只在训练集拟合避免数据泄露:# 用StandardScaler示例 scaler = preprocessing.StandardScaler().fit(X_train) X_train = scaler.transform(X_train) X_test = scaler.transform(X_test)调整训练策略
- 调整学习率:将1e-4提升至5e-4或1e-3,同时将训练轮数增加到150-200,观察损失曲线是否下降。
- 减小批次大小:把batch_size从1000改为128或256,让模型权重更新更频繁。
- 加入验证监控:训练时添加
validation_data=(X_test, Y_test),确认验证集指标是否随训练迭代提升。
排查特征有效性
- 计算特征与标签的相关性:用
data.corr()查看sales与其他特征的相关系数,如果所有特征相关性极低,说明当前特征无法区分两类,需要重新构造或选择特征。 - 删除冗余特征:移除常数、近似常数或完全无区分度的特征,减少模型噪声。
- 计算特征与标签的相关性:用
调整模型细节
- 更换输出层初始化:将最后一层的
he_normal()换成glorot_uniform(),更适合sigmoid输出层;或直接使用默认初始化。 - 尝试类别权重:即使数据集均衡,也可加入
class_weight='balanced'参数,帮助模型关注易忽略的类别。
- 更换输出层初始化:将最后一层的
验证数据处理正确性
- 检查维度匹配:打印
X.shape和Y.shape,确保输入特征数与模型input_dim一致,标签维度符合要求。 - 确认训练集分布:用
np.bincount(Y_train.flatten())检查训练集两类样本占比是否接近50%,避免分层抽样失效。
- 检查维度匹配:打印
内容的提问来源于stack exchange,提问作者Syrus
相关产品推荐
相关产品推荐

