You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas按ID随机抽取指定数量Case ID的技术问询

按ID随机抽取Case ID的解决方案

现有DataFrame

1. 含ID及工作地点数量的DataFrame

代码定义:

import pandas as pd
id_df = pd.DataFrame([[1,3],[2,2]], columns = ['ID','# of places worked'])

对应表格:

ID# of places worked
13
22

2. 含ID及Case ID的DataFrame

代码定义:

cases = pd.DataFrame(
[[1,123],[1,345],[1,456],[1,789],[1,132],[2,133],[2,143],[2,465],[2,765]], 
columns = ['ID','Case ID'])

对应表格:

IDCase ID
1123
1345
1456
1789
1132
2133
2143
2465
2765

需求

对id_df中的所有ID,从cases表中随机抽取3个对应的Case ID,期望输出示例:

IDCase ID
1456
1789
1132
2143
2465
2765

解决方案

直接用Pandas的groupby结合sample方法就能实现,完整代码如下:

import pandas as pd

# 定义两个DataFrame
id_df = pd.DataFrame([[1,3],[2,2]], columns = ['ID','# of places worked'])
cases = pd.DataFrame(
[[1,123],[1,345],[1,456],[1,789],[1,132],[2,133],[2,143],[2,465],[2,765]], 
columns = ['ID','Case ID'])

# 按ID分组,每组随机抽取3条数据
result = cases.groupby('ID').sample(n=3, random_state=42)
# 重置索引(可选,让输出索引更整洁)
result = result.reset_index(drop=True)

print(result)

代码说明

  • groupby('ID'):将cases按ID分组,每个组对应同一个ID的所有Case ID
  • sample(n=3):对每个组随机抽取3条样本;random_state=42用于固定随机种子,保证每次运行结果一致,不需要固定结果可删除该参数
  • 运行后得到的result即为符合需求的DataFrame,格式与示例一致

内容的提问来源于stack exchange,提问作者CowboyCoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 02:01:10