You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas基于id和outcome条件拆分大DataFrame为子DataFrame

基于指定条件拆分DataFrame为多个子DataFrame

问题描述

现有一个大型DataFrame,需要按以下规则拆分:找到所有id=16且outcome=1的行,向上提取数据直到遇到outcome≠1的行,最终得到多个子DataFrame。以下是示例数据和期望拆分结果:

示例原DataFrame片段:

|          | id             |outcome|
| -------- | -------------- |       | 
|          |1               |   0   |
|    0     |    1           |   1   |
|          |     2          |   1   |
|    0     |     16         |  1    |
|          |      3         |  1    |
|    0     |      5         |   1   |
|          |      4         |  8    |
|    0     |     1          |  1    |
|          |    1           |   1   |
|          |     1          |   1   |
|          |     1          |  1    |
|          |     1          |  1    |
|          |     16         | 1     |

期望拆分结果:
第一个子DataFrame:

|    0     |    1            |   1    |
|          |     2           |   1    |
|    0     |     16          |  1     |

第二个子DataFrame:

|    0     |     1           |  1     |
|          |    1            |    1   |
|          |     1           |   1    |
|          |     1           |  1     |
|          |     1           |  1     |
|          |     16          | 1      |

解决方案(Pandas实现)

直接用Pandas的索引定位和布尔筛选就能实现,代码如下:

import pandas as pd

# 构造示例DataFrame(对应你提供的片段)
data = {
    'id': [1, 1, 2, 16, 3, 5, 4, 1, 1, 1, 1, 1, 16],
    'outcome': [0, 1, 1, 1, 1, 1, 8, 1, 1, 1, 1, 1, 1]
}
df = pd.DataFrame(data)

# 1. 定位所有符合条件的目标行(id=16且outcome=1)的索引
target_indices = df[(df['id'] == 16) & (df['outcome'] == 1)].index.tolist()

# 2. 遍历每个目标索引,提取对应的子DataFrame
sub_dfs = []
for idx in target_indices:
    # 从当前目标索引向上找第一个outcome不等于1的行的索引
    first_non_one_idx = df.loc[:idx, 'outcome'].ne(1).idxmax()
    # 判断是否从开头到目标索引全是outcome=1的情况
    if df.loc[first_non_one_idx, 'outcome'] == 1:
        start_idx = 0
    else:
        start_idx = first_non_one_idx + 1
    # 提取区间内的子DataFrame并加入列表
    sub_df = df.loc[start_idx:idx, :]
    sub_dfs.append(sub_df)

# 查看拆分后的结果
for i, sub_df in enumerate(sub_dfs, 1):
    print(f"第{i}个子DataFrame:")
    print(sub_df)
    print("-" * 20)

代码说明

  • df.loc[:idx, 'outcome'].ne(1).idxmax():从DataFrame开头到当前目标索引,找到第一个outcome≠1的行的索引,ne(1)生成布尔值序列,idxmax()返回第一个True的位置。
  • 边界处理:如果从开头到目标索引的所有行outcome都是1,idxmax()会返回0,但此时该行的outcome为1,所以直接将起始索引设为0。
  • 最终所有子DataFrame会被存入sub_dfs列表,可按需调用。

内容的提问来源于stack exchange,提问作者WilliamAshoti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 10:35:55