如何不打乱数据顺序实现分层train_test_split?解决ValueError报错
解决无shuffle时的分层时间序列拆分问题
你遇到的报错是因为sklearn的train_test_split不支持在关闭打乱(shuffle=False)的同时使用分层(stratify)参数——分层逻辑依赖打乱来保证训练/测试集的类别分布一致,而时间序列要求保持原始顺序,两者在原生方法里无法兼容。
要实现指定比例拆分、保持数据顺序、同时分层的需求,需要手动实现分层采样逻辑,核心思路是:按目标变量分组,在每个类别中按比例抽取尾部样本(符合时间序列用后期数据测试的逻辑),再合并测试集,剩下的作为训练集。
实现代码
import pandas as pd # 示例数据集 data = pd.DataFrame({ 'feature1': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], 'target': [0, 1, 0, 1, 0, 1, 0, 1, 0, 1] }) test_size = 0.2 # 计算每个类别需要抽取的测试样本数(按比例取整) test_counts = (data['target'].value_counts() * test_size).round().astype(int) # 收集测试集的索引 test_indices = [] for target, count in test_counts.items(): # 取当前类别最后count个样本的索引(保持时间顺序) class_indices = data[data['target'] == target].index[-count:] test_indices.extend(class_indices) # 拆分数据集 test_df = data.loc[test_indices] train_df = data.drop(test_indices) # 按索引排序,确保训练/测试集保持原始数据顺序 train_df = train_df.sort_index() test_df = test_df.sort_index() print("训练集:") print(train_df) print("\n测试集:") print(test_df)
逻辑说明
- 先根据
test_size计算每个目标类别需要分配到测试集的样本数; - 对每个类别,选取该类别在原始数据中最后N个样本(时间序列场景下,用后期数据做测试更合理);
- 通过索引拆分出训练集和测试集,最后按索引排序保证顺序不变。
补充提示
如果test_size计算出的样本数不是整数,可根据实际需求调整取整方式(比如用math.ceil或math.floor);如果你的时间序列数据的类别分布在时间上有偏移,也可以调整采样逻辑(比如按时间窗口分层)。
内容的提问来源于stack exchange,提问作者shaik moeed
相关产品推荐
相关产品推荐

