You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何带Sigmoid输出层的分类模型预测全为NaN而非0-1值?

问题排查与解决方案

以下是针对你的二分类模型预测全为NaN的核心排查方向和解决方法:

1. 优先检查数据集的缺失值

训练或预测数据中存在未处理的NaN是导致输出全NaN的最常见原因:

  • 先统计各列缺失值情况:
    print(df.isna().sum())
    
  • 统一处理缺失值(训练集和测试集逻辑必须一致):
    # 数值型特征用中位数填充(比均值更抗异常值)
    numeric_cols = df.select_dtypes(include=['float64', 'int64']).columns
    df[numeric_cols] = df[numeric_cols].fillna(df[numeric_cols].median())
    
    # 或者直接删除含缺失值的行(数据量足够时可选)
    df = df.dropna(axis=0, how='any')
    

2. 修正特征缩放逻辑

特征值差异过大(比如部分特征取值为0-1,部分为1000+)会引发梯度爆炸,导致模型参数变为NaN:

  • 对所有数值型特征做标准化处理,注意测试集只能用训练集拟合的scaler转换:
    from sklearn.preprocessing import StandardScaler
    
    scaler = StandardScaler()
    X_train_scaled = scaler.fit_transform(X_train)
    X_test_scaled = scaler.transform(X_test)  # 测试集不能用fit_transform
    

3. 匹配模型输出层与损失函数

二分类任务用Sigmoid输出时,必须搭配BinaryCrossentropy损失函数,且from_logits参数要设为False(因为已经用了Sigmoid激活):

model.compile(
    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-4),  # 降低学习率避免梯度爆炸
    loss=tf.keras.losses.BinaryCrossentropy(from_logits=False),
    metrics=['accuracy']
)

如果误用CategoricalCrossentropy,会直接导致梯度计算异常,输出NaN。

4. 监控训练过程的异常

训练时如果loss直接变成NaN,说明训练阶段就出现了问题,可以加入早停回调终止异常训练:

early_stop = tf.keras.callbacks.EarlyStopping(
    monitor='val_loss',
    patience=3,
    restore_best_weights=True  # 恢复到loss正常的最优权重
)
history = model.fit(
    X_train_scaled, y_train,
    validation_split=0.2,
    epochs=50,
    batch_size=32,
    callbacks=[early_stop]
)

5. 确保预测数据与训练数据匹配

  • 检查预测数据的特征数量、维度是否和训练集一致(比如训练集输入是(None, 9),预测时不能是(None, 8))
  • 预测数据必须经过和训练集完全相同的预处理(缩放、编码等),不能跳过或更改逻辑

修正后的完整代码示例

import pandas as pd
import tensorflow as tf
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

# 加载Kaggle水质数据集
df = pd.read_csv('water_potability.csv')

# 处理缺失值
numeric_cols = df.select_dtypes(include=['float64', 'int64']).columns
df[numeric_cols] = df[numeric_cols].fillna(df[numeric_cols].median())

# 分离特征与标签(假设标签列为'Potability')
X = df.drop('Potability', axis=1)
y = df['Potability']

# 划分训练测试集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# 特征缩放
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

# 构建模型
model = tf.keras.Sequential([
    tf.keras.layers.Dense(32, activation='relu', input_shape=(X_train_scaled.shape[1],)),
    tf.keras.layers.Dense(16, activation='relu'),
    tf.keras.layers.Dense(1, activation='sigmoid')
])

# 编译模型
model.compile(
    optimizer=tf.keras.optimizers.Adam(learning_rate=1e-4),
    loss=tf.keras.losses.BinaryCrossentropy(from_logits=False),
    metrics=['accuracy']
)

# 训练模型
early_stop = tf.keras.callbacks.EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True)
history = model.fit(X_train_scaled, y_train, validation_split=0.2, epochs=50, batch_size=32, callbacks=[early_stop])

# 预测
predictions = model.predict(X_test_scaled)
print(predictions[:5])  # 输出0-1区间的概率值

内容的提问来源于stack exchange,提问作者JellyCZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 14:12:36