Python中如何从Pandas DataFrame的每个分类中各选取1行?
最优实现方案
直接使用pandas内置的drop_duplicates方法即可完成需求,为原生向量化实现,性能远高于循环遍历方案:
# 每个author保留第一行,再从中随机采样1000条 res = df.drop_duplicates(subset="author", keep="first").sample(n=1000, replace=False)
如果需要每个author随机保留1条而非固定取第一条,可以先打乱全表再去重:
res = df.sample(frac=1, random_state=42).drop_duplicates(subset="author", keep="first").sample(n=1000, replace=False)
备选方案(适合需对分组做额外处理的场景)
用groupby实现同等效果:
# 取每个author的第一行后采样 res = df.groupby("author", as_index=False).first().sample(n=1000, replace=False) # 每个author随机取一行后采样 res = df.groupby("author", as_index=False).apply(lambda x: x.sample(1, random_state=42)).reset_index(drop=True).sample(n=1000, replace=False)
注意事项
如果全局唯一的author数量小于1000,采样时需要将replace参数设为True开启有放回采样,否则会触发取值错误。
内容的提问来源于stack exchange,提问作者yes
相关产品推荐
相关产品推荐

