基于两列条件筛选Pandas DataFrame行并生成子DataFrame
问题需求
给定一个Pandas DataFrame,需按以下规则生成子DataFrame:当某行的typeId == 15时,提取该行之前**连续满足typeId == 1且result == 1**的行,连同该行本身存入子DataFrame。
原始DataFrame
| 索引 | typeId | result |
|---|---|---|
| 1 | 2 | 3 |
| 2 | 4 | 1 |
| 3 | 1 | 1 |
| 4 | 1 | 1 |
| 5 | 15 | 1 |
| 6 | 3 | 4 |
| 7 | 2 | 1 |
| 8 | 1 | 1 |
| 9 | 1 | 1 |
| 10 | 15 | 1 |
| 11 | 4 | 4 |
| 12 | 3 | 3 |
期望输出
第一个子DataFrame
| 索引 | typeId | result |
|---|---|---|
| 3 | 1 | 1 |
| 4 | 1 | 1 |
| 5 | 15 | 1 |
第二个子DataFrame
| 索引 | typeId | result |
|---|---|---|
| 8 | 1 | 1 |
| 9 | 1 | 1 |
| 10 | 15 | 1 |
解决方案
代码实现
import pandas as pd # 构造原始DataFrame df = pd.DataFrame({ 'typeId': [2,4,1,1,15,3,2,1,1,15,4,3], 'result': [3,1,1,1,1,4,1,1,1,1,4,3] }, index=range(1,13)) # 获取所有typeId=15的行索引 target_indices = df[df['typeId'] == 15].index.tolist() sub_dfs = [] for idx in target_indices: # 从当前索引向前遍历,收集符合条件的行 selected_rows = [idx] current = idx - 1 while current >= 1: row = df.loc[current] if row['typeId'] == 1 and row['result'] == 1: selected_rows.append(current) current -= 1 else: break # 反转索引顺序,保证与原始DataFrame一致 selected_rows.reverse() sub_df = df.loc[selected_rows] sub_dfs.append(sub_df) # 查看结果 print("第一个子DataFrame:") print(sub_dfs[0]) print("\n第二个子DataFrame:") print(sub_dfs[1])
逻辑说明
- 先定位所有
typeId=15的行索引,逐个处理 - 对每个目标行,向前遍历直到遇到不满足
typeId=1且result=1的行停止 - 反转收集的索引列表,保证子DataFrame的行顺序与原始数据一致
- 最终
sub_dfs列表包含所有符合要求的子DataFrame
内容的提问来源于stack exchange,提问作者WilliamAshoti
相关产品推荐
相关产品推荐

