如何按指定标签比例拆分多标签分类数据集,实现测试集80%class_2、15%class_1、5%class_0的占比要求
自定义测试集标签比例的数据集拆分方案
要实现测试集按80%(class_2)、15%(class_1)、5%(class_0)的比例拆分,你需要手动按标签分组抽取样本,而不是依赖train_test_split的stratify参数(它只能保持原数据集的标签比例)。下面是具体的实现步骤和代码:
步骤说明
- 先按标签将数据集分组,分离出每个类别的样本
- 计算测试集的总样本数(保持原计划的20%占比)
- 按照目标比例分配每个类别在测试集中的样本数量
- 从每个类别中随机抽取对应数量的样本作为测试集,剩余样本作为训练集
- 合并训练集和测试集的特征与标签
代码实现
假设你的数据集是一个名为df的pandas DataFrame(包含特征列和Label列):
import pandas as pd import numpy as np # 1. 按标签分组 grouped = df.groupby('Label') # 2. 计算测试集总样本数(保持原20%的比例) total_test_samples = int(len(df) * 0.2) # 3. 按目标比例分配每个类别的测试样本数 test_counts = { 2: int(total_test_samples * 0.8), 1: int(total_test_samples * 0.15), 0: int(total_test_samples * 0.05) } # 4. 抽取测试集和训练集 test_dfs = [] train_dfs = [] for label, count in test_counts.items(): # 获取当前类别的所有样本 class_df = grouped.get_group(label) # 随机抽取测试样本 test_sample = class_df.sample(n=count, random_state=42) # 剩余样本作为训练集 train_sample = class_df.drop(test_sample.index) test_dfs.append(test_sample) train_dfs.append(train_sample) # 5. 合并测试集和训练集(打乱顺序保证随机性) test_df = pd.concat(test_dfs).sample(frac=1, random_state=42) train_df = pd.concat(train_dfs).sample(frac=1, random_state=42) # 拆分特征和标签 X_train = train_df.drop('Label', axis=1) y_train = train_df['Label'] X_test = test_df.drop('Label', axis=1) y_test = test_df['Label']
验证比例
你可以用以下代码验证测试集的标签比例是否符合要求:
print("测试集标签分布:") print(y_test.value_counts(normalize=True))
注意事项
- 如果某个类别的样本数量不足以分配目标测试样本数(比如class_0的总样本数少于
total_test_samples*0.05),需要调整比例或处理样本不足的情况(比如用该类的全部样本作为测试集,再从其他类别微调数量)。 - 使用
random_state=42保证每次拆分的结果一致,方便复现实验。
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

