如何为DataFrame添加HISTORY列:识别成员组历史交集
批量标记DataFrame中ID的历史组交集
问题说明
现有一个记录成员组归属的DataFrame,包含ID(个体ID)和GROUP_NUM(组ID)两列。需求是:对每个组内的每个ID,检查该ID此前加入过的所有组,是否与当前组内其他ID此前加入过的组存在交集。若存在,在新列HISTORY标记1,否则标记0。
示例输入
| ID | GROUP_NUM |
|---|---|
| abc | 1 |
| def | 1 |
| ghi | 1 |
| jkl | 1 |
| abc | 2 |
| mno | 2 |
| pqr | 2 |
| stv | 2 |
| abc | 3 |
| stv | 3 |
| wxy | 3 |
| zzz | 3 |
| abc | 4 |
| def | 4 |
| pqr | 4 |
| bbb | 4 |
预期输出
| ID | GROUP_NUM | HISTORY |
|---|---|---|
| abc | 1 | 0 |
| def | 1 | 0 |
| ghi | 1 | 0 |
| jkl | 1 | 0 |
| abc | 2 | 0 |
| mno | 2 | 0 |
| pqr | 2 | 0 |
| stv | 2 | 0 |
| abc | 3 | 1 |
| stv | 3 | 1 |
| wxy | 3 | 0 |
| zzz | 3 | 0 |
| abc | 4 | 1 |
| def | 4 | 1 |
| pqr | 4 | 1 |
| bbb | 4 | 0 |
实现代码(Python Pandas)
import pandas as pd # 1. 构造示例数据(实际使用时替换为你的DataFrame) data = { 'ID': ['abc', 'def', 'ghi', 'jkl', 'abc', 'mno', 'pqr', 'stv', 'abc', 'stv', 'wxy', 'zzz', 'abc', 'def', 'pqr', 'bbb'], 'GROUP_NUM': [1,1,1,1,2,2,2,2,3,3,3,3,4,4,4,4] } df = pd.DataFrame(data) # 2. 生成每个ID截至当前行的历史组列表(不含当前组) df['prev_groups'] = df.groupby('ID')['GROUP_NUM'].transform( lambda x: x.shift().expanding().apply(lambda y: list(y.dropna()), raw=False) ) df['prev_groups'] = df['prev_groups'].apply(lambda x: x if isinstance(x, list) else []) # 3. 定义组内检查逻辑 def check_group_history(group): # 把组内每个ID的历史组转为集合,方便交集计算 id_history = {row['ID']: set(row['prev_groups']) for _, row in group.iterrows()} # 遍历组内每个ID,检查与其他ID历史组的交集 for idx, row in group.iterrows(): current_id = row['ID'] current_history = id_history[current_id] # 收集组内其他所有ID的历史组的并集 others_history = set() for id_, hist in id_history.items(): if id_ != current_id: others_history.update(hist) # 标记结果:有交集则为1,否则为0 group.loc[idx, 'HISTORY'] = 1 if current_history & others_history else 0 return group # 4. 按组批量处理 df = df.groupby('GROUP_NUM').apply(check_group_history) # 5. 清理临时列,得到最终结果 df = df.drop('prev_groups', axis=1) # 查看结果 print(df)
代码逻辑说明
- 生成历史组列表:通过
groupby('ID')对每个ID的组记录做偏移(shift())和累积扩展(expanding()),得到每个ID到当前行之前加入过的所有组。 - 组内交集检查:对每个组,先把所有ID的历史组转为集合(集合操作更高效),然后对每个ID,计算组内其他ID历史组的并集,再判断自己的历史组是否与该并集有交集。
- 批量处理:通过
groupby('GROUP_NUM').apply()实现对所有组的批量处理,避免逐行遍历的低效。
内容的提问来源于stack exchange,提问作者iwuzborn
相关产品推荐
相关产品推荐

