You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何从Pandas DataFrame的每个分类中各选取1行?

最优实现方案

直接使用pandas内置的drop_duplicates方法即可完成需求,为原生向量化实现,性能远高于循环遍历方案:

# 每个author保留第一行,再从中随机采样1000条
res = df.drop_duplicates(subset="author", keep="first").sample(n=1000, replace=False)

如果需要每个author随机保留1条而非固定取第一条,可以先打乱全表再去重:

res = df.sample(frac=1, random_state=42).drop_duplicates(subset="author", keep="first").sample(n=1000, replace=False)

备选方案(适合需对分组做额外处理的场景)

用groupby实现同等效果:

# 取每个author的第一行后采样
res = df.groupby("author", as_index=False).first().sample(n=1000, replace=False)

# 每个author随机取一行后采样
res = df.groupby("author", as_index=False).apply(lambda x: x.sample(1, random_state=42)).reset_index(drop=True).sample(n=1000, replace=False)

注意事项

如果全局唯一的author数量小于1000,采样时需要将replace参数设为True开启有放回采样,否则会触发取值错误。

内容的提问来源于stack exchange,提问作者yes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 10:45:05