使用Pandas采样信用卡欺诈数据集时的类别均衡问题
Hey there! I totally get the frustration—trying to grab a balanced 520-sample split between non-fraud (Class 0) and fraud (Class 1) cases from the notoriously imbalanced credit card dataset doesn’t work with basic random sampling. Let’s break down why, and fix it step by step.
Why Your Current Code Fails
The credit card fraud dataset is extremely skewed: Class 1 (fraudulent transactions) typically makes up less than 0.1% of the total data. When you run vsample_data = credit_card.sample(n=520, replace='False'), you’re just randomly picking rows—statistically, you’ll almost never get enough Class 1 samples to reach a 50/50 split. Also, quick note: replace should be a boolean (False) not a string ('False')—the string version gets treated as True, which isn’t your main issue here, but it’s good to fix that too.
Solutions to Get a Balanced Sample
Here are the most reliable ways to pull an evenly split sample:
1. Stratified Sampling (Best for Exact Control)
Use pandas to group by the Class column and sample an equal number of rows from each group. This guarantees you’ll get exactly 260 samples from Class 0 and 260 from Class 1:
# Sample 260 rows from each class to make 520 total balanced samples vsample_data = credit_card.groupby('Class', group_keys=False).apply(lambda x: x.sample(n=260))
If your dataset has fewer than 260 Class 1 samples (which is common), you can use replace=True for the minority class to oversample it:
# Oversample the minority class if there aren't enough rows vsample_data = credit_card.groupby('Class', group_keys=False).apply(lambda x: x.sample(n=260, replace=x.shape[0] < 260))
2. Use Scikit-Learn's StratifiedShuffleSplit
If you prefer using scikit-learn tools, this method also ensures proportional (or equal) sampling across classes:
from sklearn.model_selection import StratifiedShuffleSplit import pandas as pd # Set up split to get 520 samples with equal class distribution split = StratifiedShuffleSplit(n_splits=1, test_size=520, random_state=42) for train_idx, sample_idx in split.split(credit_card, credit_card['Class']): vsample_data = credit_card.iloc[sample_idx]
Note: Adjust test_size to 520 since we want the sample to be that size, not the test set.
3. Undersample the Majority Class
If you don’t mind discarding most Class 0 samples, you can undersample them to match the number of Class 1 samples, then combine with all Class 1 samples (and optionally add more via oversampling if needed):
import pandas as pd # Get all fraud samples fraud_samples = credit_card[credit_card['Class'] == 1] # Undersample non-fraud to 260 samples non_fraud_samples = credit_card[credit_card['Class'] == 0].sample(n=260, random_state=42) # Combine and shuffle the balanced sample vsample_data = pd.concat([fraud_samples, non_fraud_samples]).sample(frac=1, random_state=42)
Verify the Balance
After sampling, double-check the class distribution to make sure it’s even:
print(vsample_data['Class'].value_counts())
内容的提问来源于stack exchange,提问作者Dhruv Bhardwaj

