You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定标签比例拆分多标签分类数据集,实现测试集80%class_2、15%class_1、5%class_0的占比要求

自定义测试集标签比例的数据集拆分方案

要实现测试集按80%(class_2)、15%(class_1)、5%(class_0)的比例拆分,你需要手动按标签分组抽取样本,而不是依赖train_test_split的stratify参数(它只能保持原数据集的标签比例)。下面是具体的实现步骤和代码:

步骤说明

  1. 先按标签将数据集分组,分离出每个类别的样本
  2. 计算测试集的总样本数(保持原计划的20%占比)
  3. 按照目标比例分配每个类别在测试集中的样本数量
  4. 从每个类别中随机抽取对应数量的样本作为测试集,剩余样本作为训练集
  5. 合并训练集和测试集的特征与标签

代码实现

假设你的数据集是一个名为df的pandas DataFrame(包含特征列和Label列):

import pandas as pd
import numpy as np

# 1. 按标签分组
grouped = df.groupby('Label')

# 2. 计算测试集总样本数(保持原20%的比例)
total_test_samples = int(len(df) * 0.2)

# 3. 按目标比例分配每个类别的测试样本数
test_counts = {
    2: int(total_test_samples * 0.8),
    1: int(total_test_samples * 0.15),
    0: int(total_test_samples * 0.05)
}

# 4. 抽取测试集和训练集
test_dfs = []
train_dfs = []
for label, count in test_counts.items():
    # 获取当前类别的所有样本
    class_df = grouped.get_group(label)
    # 随机抽取测试样本
    test_sample = class_df.sample(n=count, random_state=42)
    # 剩余样本作为训练集
    train_sample = class_df.drop(test_sample.index)
    test_dfs.append(test_sample)
    train_dfs.append(train_sample)

# 5. 合并测试集和训练集(打乱顺序保证随机性)
test_df = pd.concat(test_dfs).sample(frac=1, random_state=42)
train_df = pd.concat(train_dfs).sample(frac=1, random_state=42)

# 拆分特征和标签
X_train = train_df.drop('Label', axis=1)
y_train = train_df['Label']
X_test = test_df.drop('Label', axis=1)
y_test = test_df['Label']

验证比例

你可以用以下代码验证测试集的标签比例是否符合要求:

print("测试集标签分布:")
print(y_test.value_counts(normalize=True))

注意事项

  • 如果某个类别的样本数量不足以分配目标测试样本数(比如class_0的总样本数少于total_test_samples*0.05),需要调整比例或处理样本不足的情况(比如用该类的全部样本作为测试集,再从其他类别微调数量)。
  • 使用random_state=42保证每次拆分的结果一致,方便复现实验。

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:42:40