You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Keras的简单神经网络实现疑问:二分类准确率过高排查

Troubleshooting Unusually High Accuracy in Your Binary Classification Neural Network

Hey there, that 99.79% accuracy is way higher than your expected 93-94%—definitely a sign something’s off, so let’s break down the most likely culprits (starting with the easiest checks first):

1. Data Leakage from Misapplied Scaling

This is the #1 suspect when accuracy blows past expectations. If you scaled your entire dataset before splitting into training and test sets, you’ve leaked test data statistics into your model. Here’s why:

  • When you run scaler.fit_transform(entire_dataset), the scaler uses mean/variance from both train and test data to normalize. When you later test on the "unseen" data, the model already has access to its statistical properties, making predictions unnaturally accurate.
  • Fix this by splitting first, then fitting the scaler only on training data:
    # Split data first
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
    # Fit scaler exclusively on training features
    scaler = StandardScaler()
    X_train_scaled = scaler.fit_transform(X_train)
    # Transform test data using the pre-fitted scaler
    X_test_scaled = scaler.transform(X_test)
    

2. Label Leakage in Features

Double-check if any of your 46 features directly encode the target label (benign/malicious) or are strongly correlated with it. For example:

  • A feature like malicious_status that’s a direct copy of your y variable, or a derived feature that’s just a rephrasing of the label.
  • Run correlation checks (e.g., Pearson correlation between each feature and y) to spot features with correlation values near 1 or -1. If you find one, that’s a free pass for the model to predict perfectly without learning anything meaningful.

3. Accidental Evaluation on Training Data

It’s a silly but common mistake: did you accidentally evaluate your model on the training set instead of the test set? If you ran something like model.evaluate(X_train_scaled, y_train) instead of using X_test_scaled, the model will perform nearly perfectly because it’s predicting on data it’s already seen.

4. Train/Test Split Issues

  • Did you forget to shuffle your data before splitting? If your dataset is sorted by class (e.g., all benign first, then malicious), a non-shuffled split could leave your test set dominated by one class, making accuracy look inflated.
  • Check for sample overlap between train and test sets. If some test samples are present in the training data, the model will already know their labels. You can verify this by comparing indices or using duplicate detection on the feature sets.

5. Extreme Class Imbalance

While less likely if you expected 93-94% accuracy, it’s worth ruling out: if your dataset is heavily skewed (e.g., 99% benign samples), a model that just predicts "benign" for everything will hit near-perfect accuracy but fail at identifying malicious cases. Check class counts with y_train.value_counts() and y_test.value_counts()—if imbalance exists, use metrics like precision, recall, or F1-score instead of raw accuracy to evaluate performance.

Start with the scaling check first—it’s the most common fix for this scenario. Work through each of these steps, and you’ll almost certainly find the root cause.

内容的提问来源于stack exchange,提问作者mg9893

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:15:45