You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

非随机划分数据集:确保稀有类别样本进入训练集的实现方法

嘿,我来帮你搞定这个数据集划分的问题!你遇到的这种有唯一稀有样本且必须确保它进训练集的情况,确实不能用简单的随机划分或者直接取前75%——咱们来一步步实现精准的划分,保证符合你的需求:

核心思路
  1. 先锁定稀有样本:把那个唯一的稀有类别样本单独拎出来,直接划入训练集,确保它不会被分到测试集里。
  2. 处理剩余普通样本:计算需要从普通样本中选取多少进入训练集(总训练集的75%份额减去1个稀有样本),然后对普通样本做随机划分,凑够训练集的比例。
  3. 合并结果:把稀有样本和选中的普通样本合并成最终训练集,剩下的普通样本作为测试集。
Pandas 实现代码(适合表格型数据集)

假设你的数据是用 Pandas DataFrame 存储的,类别列名为class,稀有类别值为1(你可以根据自己的实际情况修改):

import pandas as pd
import numpy as np

# 加载你的数据集(替换成你的数据路径)
df = pd.read_csv('your_dataset.csv')

# 1. 定位稀有样本和普通样本的索引
rare_class_value = 1
rare_idx = df[df['class'] == rare_class_value].index
normal_idx = df[df['class'] != rare_class_value].index

# 2. 计算需要从普通样本中选取的训练集数量
total_samples = len(df)
train_total_size = int(0.75 * total_samples)
train_normal_size = train_total_size - len(rare_idx)  # 这里len(rare_idx)=1

# 3. 从普通样本中随机选择训练集部分(无放回抽样)
train_normal_idx = np.random.choice(normal_idx, size=train_normal_size, replace=False)
test_normal_idx = [idx for idx in normal_idx if idx not in train_normal_idx]

# 4. 合并得到最终的训练集和测试集索引
train_idx = np.concatenate([rare_idx, train_normal_idx])
test_idx = test_normal_idx

# 5. 根据索引提取训练集和测试集
train_df = df.loc[train_idx]
test_df = df.loc[test_idx]
Numpy 实现代码(适合数组型特征数据)

如果你的数据是用 Numpy 数组存储的(比如特征矩阵X和标签数组y),可以这么做:

import numpy as np

# 假设X是特征数组,y是标签数组,稀有标签值为1
X = np.load('your_features.npy')
y = np.load('your_labels.npy')

# 1. 分离稀有样本和普通样本
rare_mask = y == 1
normal_mask = y != 1

X_rare, y_rare = X[rare_mask], y[rare_mask]
X_normal, y_normal = X[normal_mask], y[normal_mask]

# 2. 计算普通样本的训练集数量
total_samples = len(X)
train_total_size = int(0.75 * total_samples)
train_normal_size = train_total_size - len(X_rare)

# 3. 随机打乱普通样本索引并划分
normal_indices = np.arange(len(X_normal))
np.random.shuffle(normal_indices)

train_normal_indices = normal_indices[:train_normal_size]
test_normal_indices = normal_indices[train_normal_size:]

# 4. 合并得到最终的训练集和测试集
X_train = np.concatenate([X_rare, X_normal[train_normal_indices]])
y_train = np.concatenate([y_rare, y_normal[train_normal_indices]])
X_test = X_normal[test_normal_indices]
y_test = y_normal[test_normal_indices]
额外小贴士
  • 如果你希望划分结果可复现(每次运行代码得到相同的训练/测试集),可以在代码开头加上np.random.seed(42)(数字可以随便选,固定种子即可)。
  • 这种方式既保证了稀有样本100%在训练集,又让普通样本的划分保持随机性,比直接取前75%更合理——毕竟前75%可能会因为数据的排序问题导致分布偏差。

内容的提问来源于stack exchange,提问作者Ara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:04:50