You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas采样信用卡欺诈数据集时的类别均衡问题

Fixing Imbalanced Sampling for Credit Card Fraud Dataset

Hey there! I totally get the frustration—trying to grab a balanced 520-sample split between non-fraud (Class 0) and fraud (Class 1) cases from the notoriously imbalanced credit card dataset doesn’t work with basic random sampling. Let’s break down why, and fix it step by step.

Why Your Current Code Fails

The credit card fraud dataset is extremely skewed: Class 1 (fraudulent transactions) typically makes up less than 0.1% of the total data. When you run vsample_data = credit_card.sample(n=520, replace='False'), you’re just randomly picking rows—statistically, you’ll almost never get enough Class 1 samples to reach a 50/50 split. Also, quick note: replace should be a boolean (False) not a string ('False')—the string version gets treated as True, which isn’t your main issue here, but it’s good to fix that too.

Solutions to Get a Balanced Sample

Here are the most reliable ways to pull an evenly split sample:

1. Stratified Sampling (Best for Exact Control)

Use pandas to group by the Class column and sample an equal number of rows from each group. This guarantees you’ll get exactly 260 samples from Class 0 and 260 from Class 1:

# Sample 260 rows from each class to make 520 total balanced samples
vsample_data = credit_card.groupby('Class', group_keys=False).apply(lambda x: x.sample(n=260))

If your dataset has fewer than 260 Class 1 samples (which is common), you can use replace=True for the minority class to oversample it:

# Oversample the minority class if there aren't enough rows
vsample_data = credit_card.groupby('Class', group_keys=False).apply(lambda x: x.sample(n=260, replace=x.shape[0] < 260))

2. Use Scikit-Learn's StratifiedShuffleSplit

If you prefer using scikit-learn tools, this method also ensures proportional (or equal) sampling across classes:

from sklearn.model_selection import StratifiedShuffleSplit
import pandas as pd

# Set up split to get 520 samples with equal class distribution
split = StratifiedShuffleSplit(n_splits=1, test_size=520, random_state=42)
for train_idx, sample_idx in split.split(credit_card, credit_card['Class']):
    vsample_data = credit_card.iloc[sample_idx]

Note: Adjust test_size to 520 since we want the sample to be that size, not the test set.

3. Undersample the Majority Class

If you don’t mind discarding most Class 0 samples, you can undersample them to match the number of Class 1 samples, then combine with all Class 1 samples (and optionally add more via oversampling if needed):

import pandas as pd

# Get all fraud samples
fraud_samples = credit_card[credit_card['Class'] == 1]
# Undersample non-fraud to 260 samples
non_fraud_samples = credit_card[credit_card['Class'] == 0].sample(n=260, random_state=42)
# Combine and shuffle the balanced sample
vsample_data = pd.concat([fraud_samples, non_fraud_samples]).sample(frac=1, random_state=42)

Verify the Balance

After sampling, double-check the class distribution to make sure it’s even:

print(vsample_data['Class'].value_counts())

内容的提问来源于stack exchange,提问作者Dhruv Bhardwaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:27:12