You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按分组查找满足掩码条件的首行?

如何为每个分组查找满足掩码条件的首行?

问题背景

给定如下DataFrame:

import pandas as pd
df = pd.DataFrame(
    {
        'a': ['x', 'x', 'x', 'x', 'x', 'y', 'y', 'y', 'y', 'y', 'y', 'y'],
        'b': [1, 1, 1, 2, 2, 1, 1, 1, 2, 2, 2, 2],
        'c': [9, 8, 11, 13, 14, 3, 104, 106, 11, 100, 70, 7]
    }
)

预期输出

需要新增列out,结果如下:

a  b    c    out
0   x  1    9    NaN
1   x  1    8    NaN
2   x  1   11    NaN
3   x  2   13  found
4   x  2   14    NaN
5   y  1    3    NaN
6   y  1  104  found
7   y  1  106    NaN
8   y  2   11    NaN
9   y  2  100    NaN
10  y  2   70    NaN
11  y  2    7    NaN

掩码与处理规则

  • 掩码定义:mask = (df.c > 10)
  • 处理规则:
    • 按列a分组,每个分组内查找满足掩码条件的首行
    • 分组x需额外满足b == 2的条件,因此最终选中第3行

现有尝试代码

以下代码结果接近预期,但并非最优方案:

def func(g):
    mask = (g.c > 10)
    g.loc[mask.cumsum().eq(1) & mask, 'out'] = 'found'
    return g

df = df.groupby('a').apply(func)

优化方案

方案1:向量化操作定位目标行

先构造全局复合掩码,再按分组提取首行索引,性能更优:

# 构造复合掩码:基础掩码 + 分组x的额外条件
compound_mask = (df['c'] > 10) & ((df['a'] != 'x') | (df['b'] == 2))

# 按a分组,获取每个组内第一个满足条件的行索引
target_indices = df[compound_mask].groupby('a').head(1).index

# 初始化out列并赋值
df['out'] = pd.NA
df.loc[target_indices, 'out'] = 'found'

方案2:分组内自定义逻辑

适合需要为不同分组设置差异化规则的场景:

def mark_first_match(g):
    # 针对分组x添加额外条件
    if g.name == 'x':
        group_mask = (g['c'] > 10) & (g['b'] == 2)
    else:
        group_mask = (g['c'] > 10)
    
    g['out'] = pd.NA
    # 存在匹配行时标记首行
    if group_mask.any():
        first_idx = group_mask.idxmax()
        g.loc[first_idx, 'out'] = 'found'
    return g

df = df.groupby('a', group_keys=False).apply(mark_first_match)

两种方案均能得到预期结果,方案1更适合大规模数据,方案2扩展性更强。


内容的提问来源于stack exchange,提问作者AmirX

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 01:52:12