You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

神经网络模型准确率含义解析及信用卡欺诈分类模型准确率计算

Hey there! Let's break down what's happening with your MLP model, that surprisingly high 99.8% accuracy, and how to properly evaluate your fraud detection system.

First: That 99.8% Accuracy Is Misleading (Here's Why)

Credit card fraud datasets are extremely imbalanced—normal transactions make up 99%+ of the data, while fraud cases are just a tiny fraction (usually well under 1%).

The accuracy metric you're calculating with accuracy.eval({x: X_test, y: Y_test}) is the overall percentage of correct predictions (total correct / total samples). If your model just learns to predict "normal" for every single transaction, it will already have an accuracy close to 99.9%—even if it fails to catch any fraud at all. That's exactly what's happening here: your model is taking the easy route and ignoring the rare fraud cases.

How to Properly Evaluate Your Fraud Detector

Accuracy is useless for imbalanced datasets. You need to use metrics that focus on the minority class (fraud):

Key Metrics to Track

  • Confusion Matrix: Shows exactly how many fraud cases were caught (true positives), how many normal transactions were incorrectly flagged (false positives), and how many fraud cases were missed (false negatives).
  • Recall (True Positive Rate): The percentage of actual fraud cases your model correctly identifies—critical for fraud detection, since missing fraud is costly.
  • Precision: The percentage of transactions flagged as fraud that are actually fraud—important to avoid annoying customers with false alerts.
  • F1-Score: A balance of precision and recall, useful when you need to prioritize neither too heavily.
  • ROC-AUC: Measures how well your model can distinguish between fraud and normal transactions, even with imbalanced classes.

Code to Add These Metrics

Add this right after your existing accuracy print statement in the session:

# Get predictions and true labels
y_pred = tf.argmax(pred, 1).eval({x: X_test})
y_true = tf.argmax(Y_test.values, 1).eval(session=sess)

# Calculate and print key metrics
from sklearn.metrics import confusion_matrix, classification_report, roc_auc_score

# Confusion Matrix
print("\nConfusion Matrix:")
print(confusion_matrix(y_true, y_pred))

# Classification Report (precision, recall, F1)
print("\nClassification Report:")
print(classification_report(y_true, y_pred))

# ROC-AUC Score (uses predicted probabilities for fraud class)
y_pred_proba = tf.nn.softmax(pred).eval({x: X_test})[:, 1]
roc_auc = roc_auc_score(y_true, y_pred_proba)
print(f"\nROC-AUC Score: {roc_auc:.4f}")

Why Your Model Is Underperforming (And Fixes)

Your high accuracy hides poor fraud detection—here's how to fix the root issues:

1. Fix Class Imbalance

Your training set has way more normal transactions, so the model doesn't learn to recognize fraud. Try one of these:

  • Oversample the minority class: Duplicate fraud samples, or use SMOTE to generate synthetic fraud data.
  • Undersample the majority class: Randomly remove normal transactions to balance the dataset.
  • Weighted Loss: Assign a higher loss penalty for misclassifying fraud cases. Modify your cost function like this:
    # Assume fraud is class 1, normal is class 0
    class_weights = tf.constant([[1.0, 100.0]])  # 100x weight for fraud
    weighted_loss = tf.multiply(class_weights, tf.nn.softmax_cross_entropy_with_logits_v2(logits=pred, labels=y))
    cost = tf.reduce_mean(weighted_loss)
    

2. Preprocess Your Features

Your code uses raw features like Time and Amount which have huge value ranges compared to the V1-V28 features. MLP models are sensitive to unnormalized data—add feature scaling:

from sklearn.preprocessing import StandardScaler

# Scale features before splitting
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Now split into train/test sets
X_train, X_test, Y_train, Y_test = train_test_split(X_scaled, Y, test_size=0.3)

3. Improve Model Initialization

tf.random_normal can lead to unstable initial weights. Use Xavier/Glorot initialization instead for better training stability:

weights = {
    'h1': tf.Variable(tf.glorot_uniform_initializer()([n_input, n_hidden_1])),
    'h2': tf.Variable(tf.glorot_uniform_initializer()([n_hidden_1, n_hidden_2])),
    'out': tf.Variable(tf.glorot_uniform_initializer()([n_hidden_2, n_classes]))
}

4. Use a Proper Train/Validation/Test Split

Right now you're using a 70/30 train/test split, but you should add a validation set to monitor model performance during training (to avoid overfitting). Try a 60/20/20 split:

X_train, X_temp, Y_train, Y_temp = train_test_split(X_scaled, Y, test_size=0.4)
X_val, X_test, Y_val, Y_test = train_test_split(X_temp, Y_temp, test_size=0.5)

Then calculate metrics on the validation set after each epoch to track how well the model generalizes.

What "Final Accuracy" Actually Means (If You Still Want It)

If you still want to report accuracy, remember:

  • It only tells you how many total predictions were correct, not how well you're catching fraud.
  • Always calculate it on a held-out test set that the model never saw during training or validation.
  • Pair it with the metrics we discussed earlier to give a full picture of your model's performance.

内容的提问来源于stack exchange,提问作者Mohammad Fneish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:49:28