You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

信用卡欺诈数据集保分布抽样问题:简单随机抽样无法满足需求

Fixing Stratified Sampling for Imbalanced Credit Card Fraud Data

Hey there! The problem you're hitting is super common with imbalanced datasets like credit card fraud records—where Class-1 (fraudulent transactions) usually makes up just a tiny slice of the total data. Your current random sampling doesn't account for this imbalance, so you're probably ending up with a sample that either has no fraud cases at all, or a ratio that's way off from the original dataset.

Here are a few straightforward solutions to get a sample that preserves the original Class distribution:

1. Use Pandas' Built-in Stratified Sampling

This is the easiest way to maintain the exact class ratio from your original dataset. Just add the stratify parameter to your sample() call, and fix that replace parameter issue (you used a string 'False' instead of a boolean False—that was making Python treat it as a truthy value, so you were doing with-replacement sampling by accident!).

# Correct stratified sampling with proper boolean replace parameter
vsample_data = credit_card.sample(n=100, replace=False, stratify=credit_card['Class'])
# Verify the distribution matches the original
print(vsample_data['Class'].value_counts(normalize=True))

The stratify parameter tells pandas to split the sample proportionally across each category in the Class column, so your 100-sample subset will have the same fraud/non-fraud ratio as the full dataset.

2. Manually Specify Sample Sizes per Class

If you want explicit control over how many samples you take from each class (e.g., ensuring you get at least 5 fraud cases even if they're rare), you can split the dataset first and sample each class individually:

import pandas as pd

# Split the dataset into fraud and non-fraud subsets
non_fraud = credit_card[credit_card['Class'] == 0]
fraud = credit_card[credit_card['Class'] == 1]

# Sample desired numbers from each class (adjust counts as needed)
sample_non_fraud = non_fraud.sample(n=95, replace=False)
sample_fraud = fraud.sample(n=5, replace=False)

# Combine and shuffle the samples to avoid ordering bias
vsample_data = pd.concat([sample_non_fraud, sample_fraud]).sample(frac=1)

# Check the final counts
print(vsample_data['Class'].value_counts())

This is great if you need to guarantee a minimum number of fraud samples in your subset—critical for testing models that need to detect rare events.

Quick Pre-Check: Verify Original Dataset Distribution

Before sampling, it's always a good idea to confirm the class ratio in your full dataset so you know what to expect in your sample:

print("Original dataset class distribution:")
print(credit_card['Class'].value_counts(normalize=True))

A Quick Note on Replacement

If your original fraud dataset has fewer cases than the number you want to sample (e.g., only 3 fraud records but you want 5), you'll need to set replace=True for the fraud sample to allow duplicate entries. Just be aware that this introduces some redundancy.

内容的提问来源于stack exchange,提问作者Dhruv Bhardwaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:29:50