You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas分层采样时如何保留原索引以正确获取剩余未采样数据

解决pandas分层采样后无法匹配原索引筛选剩余数据的问题

问题原因

你出现该问题的核心原因是groupby().apply()默认会将分组键作为多级索引的第一层,采样得到的dfsample索引为两级结构,第一层是Sentiment的分组值,第二层才是原DataFrame的行索引,因此直接用原索引匹配时无法命中,导致剩余数据筛选失败。

解决方案

方案1:使用group_keys参数关闭分组索引(推荐)

pandas 1.1.0及以上版本支持groupby的group_keys参数,设为False后不会将分组键加入结果索引,直接保留原行索引:

import pandas as pd

# 示例数据集
df = pd.DataFrame({'text':['hello', 'how', 'good', 'bad', 'ok', 'bye', 'ol'], 'Sentiment':[0, 1, 1, 1, 2, 0, 2]})

# 分层采样:每组抽1条,要按20%比例抽取替换为frac=0.2即可
dfsample = df.groupby('Sentiment', group_keys=False).apply(lambda x: x.sample(n=1, random_state=42))

# 筛选剩余未采样数据
df_rest = df.loc[~df.index.isin(dfsample.index)]

方案2:手动删除多级索引的分组层

如果使用的pandas版本较低不支持group_keys参数,可以手动删除多级索引的第一层分组键:

dfsample = df.groupby('Sentiment').apply(lambda x: x.sample(n=1, random_state=42)).droplevel(0)
df_rest = df.loc[~df.index.isin(dfsample.index)]

结果验证

运行后得到的df_rest完全符合预期输出,采样得到的dfsample索引和原df完全一致,不会出现匹配失效的问题。

注意事项

  • 如果按比例采样时部分类别样本量过少,可按需添加replace=True参数允许有放回采样,避免报错。
  • 添加random_state参数可以固定采样结果,方便复现实验。

内容的提问来源于stack exchange,提问作者sariii

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 02:15:02