You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

实现基于目标比例的SMOTE过采样与随机欠采样函数报错排查

问题分析与解决方案

第一个函数报错原因

你错误地将单个类的子集单独传入SMOTE/欠采样器:

  • SMOTE的核心逻辑是基于同类样本与邻近异类样本的分布合成新数据,必须依赖多类别数据集才能工作。单独传入单个类的标签时,采样器会检测到只有1个类别,触发ValueError: The target 'y' needs to have more than 1 class。

第二个函数报错原因

SMOTE的sampling_strategy字典参数要求传入目标样本数,而非增量样本数:

  • 你的代码计算的oversample_amount是需要新增的样本数(比如447),远小于原始样本数1248,采样器会误认为你要将该类样本数减少到447,违背过采样逻辑,因此报错With over-sampling methods, the number of samples in a class should be greater or equal to the original number of samples。

修复后的重采样函数

以下函数先对需要欠采样的类别进行处理(避免后续过采样的样本被浪费),再用SMOTE完成过采样,同时处理目标比例归一化、边界值检查等问题:

from imblearn.over_sampling import SMOTE
from imblearn.under_sampling import RandomUnderSampler
from collections import Counter
import numpy as np

def resample_to_proportion(X, y, target_proportion, normalize_target=True):
    # 计算当前类别样本数与比例
    class_counts = Counter(y)
    total_samples = len(y)
    current_proportions = {label: count / total_samples for label, count in class_counts.items()}
    
    # 归一化目标比例(确保总和为1,可选关闭)
    if normalize_target:
        total_target_sum = sum(target_proportion.values())
        target_proportion = {label: prop / total_target_sum for label, prop in target_proportion.items()}
    
    # 计算每个类别的目标样本数(基于原始总样本数的比例)
    target_counts = {}
    for label in class_counts:
        if label not in target_proportion:
            # 未指定目标比例的类别保持原样本数
            target_counts[label] = class_counts[label]
        else:
            # 计算目标样本数并做边界校验
            target_count = round(total_samples * target_proportion[label])
            # 欠采样时至少保留1个样本
            target_counts[label] = max(1, target_count)
            # 过采样时目标数不小于原始样本数
            if target_proportion[label] > current_proportions[label]:
                target_counts[label] = max(target_counts[label], class_counts[label])
    
    # 第一步:对需要欠采样的类别进行处理
    under_strategy = {label: cnt for label, cnt in target_counts.items() if cnt < class_counts[label]}
    if under_strategy:
        undersampler = RandomUnderSampler(sampling_strategy=under_strategy, random_state=42)
        X, y = undersampler.fit_resample(X, y)
        class_counts = Counter(y)  # 更新采样后的类别计数
    
    # 第二步:对需要过采样的类别进行处理
    over_strategy = {label: cnt for label, cnt in target_counts.items() if cnt > class_counts[label]}
    if over_strategy:
        smote = SMOTE(sampling_strategy=over_strategy, random_state=42)
        X, y = smote.fit_resample(X, y)
    
    return X, y

测试示例

用你提供的场景验证函数:

from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split

# 生成测试数据
X, y = make_classification(n_samples=10000, n_features=10, n_classes=5,
                           n_informative=4, weights=[0.3,0.125,0.239,0.153,0.188], random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

# 定义目标比例
target_proportion = {0: 0.519, 1: 0.373, 2: 0.226, 3: 0.053, 4: 0.164}

# 执行重采样
X_resampled, y_resampled = resample_to_proportion(X_train, y_train, target_proportion)

# 输出重采样后的比例
resampled_counts = Counter(y_resampled)
total_resampled = len(y_resampled)
print("重采样后类别比例:")
for label in sorted(resampled_counts):
    print(f"class {label}: {resampled_counts[label]/total_resampled:.3f}")

输出结果会接近你设定的目标比例(因取整操作存在微小误差)。


内容的提问来源于stack exchange,提问作者Amina Umar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 01:19:52