实现基于目标比例的SMOTE过采样与随机欠采样函数报错排查
问题分析与解决方案
第一个函数报错原因
你错误地将单个类的子集单独传入SMOTE/欠采样器:
- SMOTE的核心逻辑是基于同类样本与邻近异类样本的分布合成新数据,必须依赖多类别数据集才能工作。单独传入单个类的标签时,采样器会检测到只有1个类别,触发
ValueError: The target 'y' needs to have more than 1 class。
第二个函数报错原因
SMOTE的sampling_strategy字典参数要求传入目标样本数,而非增量样本数:
- 你的代码计算的
oversample_amount是需要新增的样本数(比如447),远小于原始样本数1248,采样器会误认为你要将该类样本数减少到447,违背过采样逻辑,因此报错With over-sampling methods, the number of samples in a class should be greater or equal to the original number of samples。
修复后的重采样函数
以下函数先对需要欠采样的类别进行处理(避免后续过采样的样本被浪费),再用SMOTE完成过采样,同时处理目标比例归一化、边界值检查等问题:
from imblearn.over_sampling import SMOTE from imblearn.under_sampling import RandomUnderSampler from collections import Counter import numpy as np def resample_to_proportion(X, y, target_proportion, normalize_target=True): # 计算当前类别样本数与比例 class_counts = Counter(y) total_samples = len(y) current_proportions = {label: count / total_samples for label, count in class_counts.items()} # 归一化目标比例(确保总和为1,可选关闭) if normalize_target: total_target_sum = sum(target_proportion.values()) target_proportion = {label: prop / total_target_sum for label, prop in target_proportion.items()} # 计算每个类别的目标样本数(基于原始总样本数的比例) target_counts = {} for label in class_counts: if label not in target_proportion: # 未指定目标比例的类别保持原样本数 target_counts[label] = class_counts[label] else: # 计算目标样本数并做边界校验 target_count = round(total_samples * target_proportion[label]) # 欠采样时至少保留1个样本 target_counts[label] = max(1, target_count) # 过采样时目标数不小于原始样本数 if target_proportion[label] > current_proportions[label]: target_counts[label] = max(target_counts[label], class_counts[label]) # 第一步:对需要欠采样的类别进行处理 under_strategy = {label: cnt for label, cnt in target_counts.items() if cnt < class_counts[label]} if under_strategy: undersampler = RandomUnderSampler(sampling_strategy=under_strategy, random_state=42) X, y = undersampler.fit_resample(X, y) class_counts = Counter(y) # 更新采样后的类别计数 # 第二步:对需要过采样的类别进行处理 over_strategy = {label: cnt for label, cnt in target_counts.items() if cnt > class_counts[label]} if over_strategy: smote = SMOTE(sampling_strategy=over_strategy, random_state=42) X, y = smote.fit_resample(X, y) return X, y
测试示例
用你提供的场景验证函数:
from sklearn.datasets import make_classification from sklearn.model_selection import train_test_split # 生成测试数据 X, y = make_classification(n_samples=10000, n_features=10, n_classes=5, n_informative=4, weights=[0.3,0.125,0.239,0.153,0.188], random_state=42) X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42) # 定义目标比例 target_proportion = {0: 0.519, 1: 0.373, 2: 0.226, 3: 0.053, 4: 0.164} # 执行重采样 X_resampled, y_resampled = resample_to_proportion(X_train, y_train, target_proportion) # 输出重采样后的比例 resampled_counts = Counter(y_resampled) total_resampled = len(y_resampled) print("重采样后类别比例:") for label in sorted(resampled_counts): print(f"class {label}: {resampled_counts[label]/total_resampled:.3f}")
输出结果会接近你设定的目标比例(因取整操作存在微小误差)。
内容的提问来源于stack exchange,提问作者Amina Umar
相关产品推荐
相关产品推荐

