基于特定架构的TensorFlow二分类模型训练准确率持续为0的排查求助
Hey there, that's a tricky situation—your loss is plummeting all the way to near-zero, but accuracy stays stuck at 0%? That tells me the model is "learning" to minimize loss, but it's completely missing the mark when it comes to classifying your data. Let's break down the most likely issues and fixes for your setup (5 features, binary target vmcategory where 0 = "Delay-insensitive" and 1 = "Interactive"):
1. Mismatched Loss Function & Output Layer
This is the #1 culprit for this exact symptom. For binary classification:
- If your output layer uses
sigmoidactivation (the standard choice for 0/1 labels), you must usebinary_crossentropyas your loss function. Usingcategorical_crossentropyhere will break accuracy calculations because it expects one-hot encoded labels, not raw 0/1 values. - Double-check your model compile code—it should look like this:
model.compile( optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'] )
Also, make sure your output layer has sigmoid activation:
model.add(tf.keras.layers.Dense(1, activation='sigmoid'))
Without that sigmoid, your model outputs unbounded values, and the accuracy metric can't properly map them to 0/1 classes.
2. Data Preprocessing Red Flags
Severe Class Imbalance
First, check how many samples you have for each class. If 99% of your data is one class, the model might be predicting that class every time—but since accuracy is 0%, it's probably the opposite: the model is always predicting the minority class, and your training set is almost entirely the majority class.
- Verify with:
print(df['vmcategory'].value_counts())
Fix this by either:
- Resampling (oversample the minority class or undersample the majority)
- Using
class_weightduring training to give more weight to the minority class:
class_weights = { 0: len(df[df['vmcategory'] == 1]) / len(df), 1: len(df[df['vmcategory'] == 0]) / len(df) } model.fit(X_train, y_train, class_weight=class_weights, epochs=100)
Unnormalized Features
If your 5 features have wildly different value ranges (e.g., one is 0-1, another is 1000-10000), the model will struggle to learn meaningful patterns, often leading to biased predictions.
- Normalize your features with
StandardScalerorMinMaxScaler:
from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_train_scaled = scaler.fit_transform(X_train) X_test_scaled = scaler.transform(X_test)
Use the scaled features for training instead of raw values.
3. Odd Training Setup
I noticed each epoch only runs 1 step—that means your batch_size is equal to your entire training dataset size. Training on the full batch every time can make the model's updates unstable and slow to learn generalizable patterns.
- Try reducing your batch size to a standard value like 32 or 64:
model.fit(X_train_scaled, y_train, batch_size=32, epochs=100)
4. Verify Label Integrity
Quick sanity check: make sure your target labels are actually only 0 and 1. If there's a typo or unexpected value in vmcategory, the model can't learn correctly.
- Run:
print(np.unique(y_train))
If you see values other than 0 and 1, clean your dataset first.
5. Manual Accuracy Check
To rule out any metric calculation bugs in TensorFlow, compute accuracy manually after training:
import numpy as np y_pred_probs = model.predict(X_train_scaled) y_pred_classes = (y_pred_probs > 0.5).astype(int) # Threshold at 0.5 for binary classification manual_accuracy = np.mean(y_pred_classes == y_train) print(f"Manual Training Accuracy: {manual_accuracy:.2f}")
If this is also 0%, you know the model is truly predicting the wrong class every time—focus on the earlier fixes. If it's higher, there's a bug in how you're compiling the model's metrics (though this is rare).
Your Training Log for Reference
Epoch 1/100 1/1 [] - 29s 29s/step - loss: 0.6931 - accuracy: 0.0000e+00
Epoch 2/100 1/1 [] - 7s 7s/step - loss: 0.6893 - accuracy: 0.0000e+00
Epoch 3/100 1/1 [] - 7s 7s/step - loss: 0.6808 - accuracy: 0.0000e+00
Epoch 4/100 1/1 [] - 7s 7s/step - loss: 0.6571 - accuracy: 0.0000e+00
Epoch 5/100 1/1 [] - 7s 7s/step - loss: 0.5957 - accuracy: 0.0000e+00
Epoch 6/100 1/1 [] - 7s 7s/step - loss: 0.5372 - accuracy: 0.0000e+00
Epoch 7/100 1/1 [] - 7s 7s/step - loss: 0.3760 - accuracy: 0.0000e+00
Epoch 8/100 1/1 [] - 7s 7s/step - loss: 0.2411 - accuracy: 0.0000e+00
Epoch 9/100 1/1 [] - 7s 7s/step - loss: 0.1913 - accuracy: 0.0000e+00
Epoch 10/100 1/1 [] - 7s 7s/step - loss: 0.0571 - accuracy: 0.0000e+00
Epoch 11/100 1/1 [] - 7s 7s/step - loss: 0.0483 - accuracy: 0.0000e+00
Epoch 12/100 1/1 [] - 7s 7s/step - loss: 0.0088 - accuracy: 0.0000e+00
Epoch 13/100 1/1 [] - 7s 7s/step - loss: 6.1697e-04 - accuracy: 0.0000e+00
Epoch 14/100 1/1 [] - 6s 6s/step - loss: 3.2386e-04 - accuracy: 0.0000e+00
Epoch 15/100 1/1 [] - 6s 6s/step - loss: 6.8086e-06 - accuracy: 0.0000e+00
Epoch 16/100 1/1 [] - 6s 6s/step - loss: 7.7796e-05 - accuracy: 0.0000e+00
Epoch 17/100 1/1 [] - 7s 7s/step - loss: 1.1021e-06 - accuracy: 0.0000e+00
Epoch 18/100 1/1 [] - 6s 6s/step - loss: 2.7273e-07 - accuracy: 0.0000e+00
Epoch 87/100 1/1 [] - 6s 6s/step - loss: 1.0003e-13 - accuracy: 0.0000e+00
Epoch 88/100 1/1 [] - 6s 6s/step - loss: 2.6685e-14 - accuracy: 0.0000e+00
Epoch 89/100 1/1 [] - 7s 7s/step - loss: 2.4792e-12 - accuracy: 0.0000e+00
Epoch 90/100 1/1 [] - 7s 7s/step - loss: 1.2417e-13 - accuracy: 0.0000e+00
Epoch 91/100 1/1 [] - 7s 7s/step - loss: 1.4707e-11 - accuracy: 0.0000e+00
Epoch 92/100 1/1 [] - 7s 7s/step - loss: 4.9625e-14 - accuracy: 0.0000e+00
Epoch 93/100 1/1 [] - 7s 7s/step - loss: 3.7239e-13 - accuracy: 0.0000e+00
Epoch 94/100 1/1 [] - 7s 7s/step - loss: 6.0243e-13 - accuracy: 0.0000e+00
Epoch 95/100 1/1 [] - 6s 6s/step - loss: 1.4047e-11 - accuracy: 0.0000e+00
Epoch 96/100 1/1 [] - 7s 7s/step - loss: 1.0687e-14 - accuracy: 0.0000e+00
Epoch 97/100 1/1 [] - 7s 7s/step - loss: 3.4614e-16 - accuracy: 0.0000e+00
Epoch 98/100 1/1 [] - 7s 7s/step - loss: 4.5617e-11 - accuracy: 0.0000e+00
Epoch 99/100 1/1 [] - 7s 7s/step - loss: 1.5913e-14 - accuracy: 0.0000e+00
Epoch 100/100 1/1 [] - 7s 7s/step - loss: 3.0236e-10 - accuracy: 0.0000e+00
内容的提问来源于stack exchange,提问作者MexcelsiorB

