You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用sklearn的train_test_split按类别等比例拆分dataframe数据集

解决方案

你要的按指定分类列同比例拆分的需求,直接用sklearn.model_selection.train_test_split的stratify参数就能实现,不需要手动拆分多个数据集再合并。

方法1:直接修改原有train_test_split调用

你只需要在原有拆分代码里新增stratify参数,传入你要做分层依据的category列即可,修改后代码如下:

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split

nrows=1000
cat_values = ['A','B','C','D']
cats = np.random.choice(cat_values, size=(nrows))

df = pd.DataFrame(np.random.randint(0,1000,size=(nrows, 3)), columns=['variable 1','variable 2','variable 3'])
df['category'] = cats
y = np.random.rand(nrows)

# 新增stratify参数,指定按category列分层
X_train, X_test, y_train, y_test = train_test_split(df, y, test_size = .2, random_state =0, stratify=df['category'])

# 验证每个类别拆分比例
for cat in cat_values:
    total = len(df[df['category']==cat])
    train_cnt = len(X_train[X_train['category']==cat])
    print(f"类别{cat}:总样本数{total},训练集数{train_cnt},训练集占比{train_cnt/total:.1%}")

运行后可以看到每个类别的训练集占比都稳定在80%左右,符合你的需求。

方法2:用StratifiedShuffleSplit实现

你之前用StratifiedShuffleSplit没找到指定分层列的入口,是因为分层依据是在split()方法调用时传入的第三个参数,示例代码如下:

from sklearn.model_selection import StratifiedShuffleSplit

sss = StratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=0)
# split方法第三个参数就是分层依据,这里传入df['category']
for train_idx, test_idx in sss.split(df, y, df['category']):
    X_train, X_test = df.iloc[train_idx], df.iloc[test_idx]
    y_train, y_test = y[train_idx], y[test_idx]

这个方法和上面的train_test_split效果一致,适合需要多次重复拆分的场景。


内容的提问来源于stack exchange,提问作者Hoppity81

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 18:36:00