You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不打乱数据顺序实现分层train_test_split?解决ValueError报错

解决无shuffle时的分层时间序列拆分问题

你遇到的报错是因为sklearn的train_test_split不支持在关闭打乱(shuffle=False)的同时使用分层(stratify)参数——分层逻辑依赖打乱来保证训练/测试集的类别分布一致,而时间序列要求保持原始顺序,两者在原生方法里无法兼容。

要实现指定比例拆分、保持数据顺序、同时分层的需求,需要手动实现分层采样逻辑,核心思路是:按目标变量分组,在每个类别中按比例抽取尾部样本(符合时间序列用后期数据测试的逻辑),再合并测试集,剩下的作为训练集。

实现代码

import pandas as pd

# 示例数据集
data = pd.DataFrame({
    'feature1': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10],
    'target': [0, 1, 0, 1, 0, 1, 0, 1, 0, 1]
})

test_size = 0.2
# 计算每个类别需要抽取的测试样本数(按比例取整)
test_counts = (data['target'].value_counts() * test_size).round().astype(int)

# 收集测试集的索引
test_indices = []
for target, count in test_counts.items():
    # 取当前类别最后count个样本的索引(保持时间顺序)
    class_indices = data[data['target'] == target].index[-count:]
    test_indices.extend(class_indices)

# 拆分数据集
test_df = data.loc[test_indices]
train_df = data.drop(test_indices)

# 按索引排序,确保训练/测试集保持原始数据顺序
train_df = train_df.sort_index()
test_df = test_df.sort_index()

print("训练集:")
print(train_df)
print("\n测试集:")
print(test_df)

逻辑说明

  1. 先根据test_size计算每个目标类别需要分配到测试集的样本数;
  2. 对每个类别,选取该类别在原始数据中最后N个样本(时间序列场景下,用后期数据做测试更合理);
  3. 通过索引拆分出训练集和测试集,最后按索引排序保证顺序不变。

补充提示

如果test_size计算出的样本数不是整数,可根据实际需求调整取整方式(比如用math.ceil或math.floor);如果你的时间序列数据的类别分布在时间上有偏移,也可以调整采样逻辑(比如按时间窗口分层)。

内容的提问来源于stack exchange,提问作者shaik moeed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 13:19:51