You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于UCI肝炎数据集使用ADASYN时X、Y的取值与拆分问询

Handling Class Imbalance in UCI Hepatitis Dataset with ADASYN

Let's walk through this step by step—from splitting your dataset into features and targets to applying ADASYN properly. I'll keep this practical and code-focused, tailored to your specific dataset.

Step 1: Load and Clean the Dataset

First, load the UCI Hepatitis dataset (I assume you’ve downloaded it already). Most versions come without headers, so we’ll use pandas to read it, then handle missing values (this dataset has plenty of them, which will break ADASYN if left unaddressed).

import pandas as pd

# Load the dataset (update the file path to match your local file)
df = pd.read_csv('hepatitis.data', header=None)

# Replace '?' (common missing value marker in UCI datasets) with NaN
df = df.replace('?', pd.NA)

# Convert all columns to numeric (since some features might be stored as strings)
df = df.apply(pd.to_numeric, errors='coerce')

Step 2: Split into Features (X) and Target (Y)

The Hepatitis dataset’s last column is the target class: typically coded as 1 for LIVE and 2 for DIE (double-check your dataset docs to confirm). We’ll map these to 0 (DIE) and 1 (LIVE) for easier handling with ADASYN:

# Target (Y): last column, mapped to 0 (minority class: DIE) and 1 (majority: LIVE)
y = df.iloc[:, -1].map({2: 0, 1: 1})

# Features (X): all columns except the last one
X = df.iloc[:, :-1]

# Fill missing values in features (use mean for continuous, median for skewed columns if needed)
X = X.fillna(X.mean())

Step 3: Split into Training and Test Sets (Critical!)

Never apply oversampling to the entire dataset—this causes data leakage, which makes your model’s performance look better than it actually is. Always split into train/test first, then only oversample the training set. Use stratify=y to preserve the original class balance in both sets:

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y,
    test_size=0.2,  # Reserve 20% of data for testing
    random_state=42,  # For reproducible results
    stratify=y  # Keep class distribution consistent across train/test
)

Step 4: Apply ADASYN to the Training Set

Now we can use ADASYN to generate synthetic samples for the minority class (DIE, 32 samples) to balance the training data:

from imblearn.over_sampling import ADASYN

# Initialize ADASYN (you can tweak n_neighbors if needed—default is 5)
adasyn = ADASYN(random_state=42)

# Fit on training data and generate resampled balanced data
X_train_resampled, y_train_resampled = adasyn.fit_resample(X_train, y_train)

You can verify the new class distribution to confirm balance:

print(pd.Series(y_train_resampled).value_counts())
# Output should show roughly equal counts for 0 (DIE) and 1 (LIVE)

Key Tips

  • Preprocessing first: Always handle missing values and scale features (if using distance-based models like SVM or KNN) before applying ADASYN—it relies on calculating distances between samples.
  • Test set stays untouched: The test set must remain in its original imbalanced state to get an unbiased evaluation of how your model performs on real-world data.
  • Tweak parameters: If ADASYN isn’t balancing the classes as expected, adjust the n_neighbors parameter to control how many nearby samples are used to generate synthetic data.

内容的提问来源于stack exchange,提问作者Minu Bharatheedasan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:37:42