如何使用Pandas按color和type分组随机选取一行?
使用Pandas按指定列分组并随机保留每组一行
可以通过Pandas的groupby()结合sample()方法轻松实现这个需求,具体步骤如下:
- 构造示例DataFrame
先把你提供的表格数据转换成Pandas DataFrame:
import pandas as pd data = { 'id': [123, 456, 603, 673, 897, 578], 'color': ['red', 'red', 'red', 'blue', 'blue', 'black'], 'type': ['civic', 'civic', 'civic', 'rav4', 'rav4', 'rav4'], 'category': ['single', 'family', 'single', 'single', 'family', 'family'], 'age': [21, 35, 32, 23, 54, 63], 'location': ['california', 'michigan', 'seattle', 'toranto', 'texas', 'california'] } df = pd.DataFrame(data)
- 分组并随机选取每组一行
使用groupby()按color和type分组,再调用sample(n=1)从每组随机抽取一行:
result_df = df.groupby(['color', 'type']).sample(n=1)
如果你需要固定随机结果(保证每次运行输出一致),可以添加random_state参数,比如:
result_df = df.groupby(['color', 'type']).sample(n=1, random_state=42)
- 查看结果
执行上述代码后,result_df就是你需要的结果,示例输出(因随机性可能和你给出的示例略有不同):
| id | color | type | category | age | location | |
|---|---|---|---|---|---|---|
| 5 | 578 | black | rav4 | family | 63 | california |
| 4 | 897 | blue | rav4 | family | 54 | texas |
| 0 | 123 | red | civic | single | 21 | california |
关键说明
groupby(['color', 'type']):将数据按照color和type的组合进行分组,相同color+type的行归为一组。sample(n=1):从每个分组中随机选取1行数据,默认采用无放回抽样(此处每组仅取1行,无放回逻辑不影响结果)。
内容的提问来源于stack exchange,提问作者Danny
相关产品推荐
相关产品推荐

