You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按比例拆分数据集为训练/测试集并保证Key不跨集

解决按Key拆分训练/测试集的思路

核心需求是同一Key的所有样本只能出现在训练集或测试集中,同时整体样本量比例近似70/30。GroupKFold这类交叉验证工具无法直接满足,因为它的目标是保证组不跨fold,但不控制整体比例。以下是具体实现思路:

步骤1:统计每个Key的样本量

首先需要明确每个Key对应多少条数据,这是后续加权拆分的基础。用Pandas实现示例:

import pandas as pd

# 假设你的数据集是df,Key列名为'Key'
key_sample_counts = df.groupby('Key').size().reset_index(name='sample_count')
total_samples = df.shape[0]
target_train_samples = int(total_samples * 0.7)  # 目标训练集样本量

步骤2:加权拆分Key集合

由于不同Key的样本量差异极大,直接随机拆分Key会导致整体比例严重偏离目标。这里推荐两种可靠方法:

方法一:贪心加权选择(保证比例最接近)

将Key按样本量从大到小排序,依次加入训练集,直到总样本量接近目标值:

# 按样本量降序排序Key
key_sample_counts_sorted = key_sample_counts.sort_values('sample_count', ascending=False)

train_keys = []
current_train_size = 0

for _, row in key_sample_counts_sorted.iterrows():
    key = row['Key']
    count = row['sample_count']
    # 如果加入该Key后不超过目标值,直接加入
    if current_train_size + count <= target_train_samples:
        train_keys.append(key)
        current_train_size += count
    else:
        # 允许小范围误差(比如5%),避免因单个大Key导致比例偏差过大
        if (target_train_samples - current_train_size) / target_train_samples < 0.05:
            train_keys.append(key)
            current_train_size += count
        break

方法二:加权随机采样(兼顾随机性和比例)

如果需要一定的随机性,可按每个Key的样本量占比设置采样权重,多次采样直到满足比例:

import numpy as np

# 计算每个Key的采样权重(样本量占总样本的比例)
key_sample_counts['weight'] = key_sample_counts['sample_count'] / total_samples

train_keys = []
current_train_size = 0

while current_train_size < target_train_samples * 0.95:  # 允许5%的下限误差
    # 按权重随机选一个Key,避免重复选择
    candidate_key = np.random.choice(key_sample_counts[~key_sample_counts['Key'].isin(train_keys)]['Key'], 
                                    p=key_sample_counts[~key_sample_counts['Key'].isin(train_keys)]['weight'])
    candidate_count = key_sample_counts[key_sample_counts['Key'] == candidate_key]['sample_count'].values[0]
    # 加入后不超过上限则添加
    if current_train_size + candidate_count <= target_train_samples * 1.05:  # 允许5%的上限误差
        train_keys.append(candidate_key)
        current_train_size += candidate_count

步骤3:拆分数据集

根据选中的训练Key拆分原始数据:

train_df = df[df['Key'].isin(train_keys)]
test_df = df[~df['Key'].isin(train_keys)]

# 验证比例
print(f"训练集样本占比:{train_df.shape[0]/total_samples:.2f}")
print(f"测试集样本占比:{test_df.shape[0]/total_samples:.2f}")

注意事项

  • 如果存在单个Key样本量占比极高(比如超过30%),则无法严格保证70/30比例,只能尽量接近,此时需根据业务需求决定该Key放在训练集还是测试集。
  • 若需要多次拆分(比如做不同的实验),推荐用方法二,保证每次拆分的随机性;若需要最稳定的比例,优先用方法一。

内容的提问来源于stack exchange,提问作者Shadi Iskander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 00:11:08