如何用Pandas基于id和outcome条件拆分大DataFrame为子DataFrame
基于指定条件拆分DataFrame为多个子DataFrame
问题描述
现有一个大型DataFrame,需要按以下规则拆分:找到所有id=16且outcome=1的行,向上提取数据直到遇到outcome≠1的行,最终得到多个子DataFrame。以下是示例数据和期望拆分结果:
示例原DataFrame片段:
| | id |outcome| | -------- | -------------- | | | |1 | 0 | | 0 | 1 | 1 | | | 2 | 1 | | 0 | 16 | 1 | | | 3 | 1 | | 0 | 5 | 1 | | | 4 | 8 | | 0 | 1 | 1 | | | 1 | 1 | | | 1 | 1 | | | 1 | 1 | | | 1 | 1 | | | 16 | 1 |
期望拆分结果:
第一个子DataFrame:
| 0 | 1 | 1 | | | 2 | 1 | | 0 | 16 | 1 |
第二个子DataFrame:
| 0 | 1 | 1 | | | 1 | 1 | | | 1 | 1 | | | 1 | 1 | | | 1 | 1 | | | 16 | 1 |
解决方案(Pandas实现)
直接用Pandas的索引定位和布尔筛选就能实现,代码如下:
import pandas as pd # 构造示例DataFrame(对应你提供的片段) data = { 'id': [1, 1, 2, 16, 3, 5, 4, 1, 1, 1, 1, 1, 16], 'outcome': [0, 1, 1, 1, 1, 1, 8, 1, 1, 1, 1, 1, 1] } df = pd.DataFrame(data) # 1. 定位所有符合条件的目标行(id=16且outcome=1)的索引 target_indices = df[(df['id'] == 16) & (df['outcome'] == 1)].index.tolist() # 2. 遍历每个目标索引,提取对应的子DataFrame sub_dfs = [] for idx in target_indices: # 从当前目标索引向上找第一个outcome不等于1的行的索引 first_non_one_idx = df.loc[:idx, 'outcome'].ne(1).idxmax() # 判断是否从开头到目标索引全是outcome=1的情况 if df.loc[first_non_one_idx, 'outcome'] == 1: start_idx = 0 else: start_idx = first_non_one_idx + 1 # 提取区间内的子DataFrame并加入列表 sub_df = df.loc[start_idx:idx, :] sub_dfs.append(sub_df) # 查看拆分后的结果 for i, sub_df in enumerate(sub_dfs, 1): print(f"第{i}个子DataFrame:") print(sub_df) print("-" * 20)
代码说明
df.loc[:idx, 'outcome'].ne(1).idxmax():从DataFrame开头到当前目标索引,找到第一个outcome≠1的行的索引,ne(1)生成布尔值序列,idxmax()返回第一个True的位置。- 边界处理:如果从开头到目标索引的所有行
outcome都是1,idxmax()会返回0,但此时该行的outcome为1,所以直接将起始索引设为0。 - 最终所有子DataFrame会被存入
sub_dfs列表,可按需调用。
内容的提问来源于stack exchange,提问作者WilliamAshoti
相关产品推荐
相关产品推荐

