You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何随机选取DataFrame行并允许重复行?

解决Pandas分组后允许重复行的随机采样问题

嘿,这个需求我之前也碰到过!你当前用的sample(frac=...)默认是无放回采样(replace=False),所以没法选中重复的行。要实现有放回的随机选取,只需要给sample()方法加上replace=True参数就搞定了。

修改你的现有代码

直接在sample调用里添加replace=True参数,就能开启有放回采样:

dfSub = dataFrame.groupby(self.column, as_index=False).apply(
    lambda x: x.sample(frac=self.percentage, replace=True)
).reset_index(drop=True)

原理说明

当replace=True时,Pandas会在每次采样后把选中的行“放回”原分组,这样同一行就有机会被多次选中,完美满足你需要重复行的需求。

另一种灵活实现方式(用NumPy生成索引)

如果你想更精细地控制采样逻辑,也可以用numpy.random.choice生成随机索引,再通过iloc提取行:

import numpy as np

dfSub = dataFrame.groupby(self.column, as_index=False).apply(
    lambda x: x.iloc[np.random.choice(
        len(x), 
        size=int(len(x)*self.percentage), 
        replace=True
    )]
).reset_index(drop=True)

这种方式和sample(replace=True)效果完全一致,适合需要自定义采样权重或其他特殊逻辑的场景。

内容的提问来源于stack exchange,提问作者Mateus Jose

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:41:56