You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按组依据AAA与BBB的出现顺序设置DataFrame的TGT列值?

问题描述

现有如下DataFrame:

group  AAA  BBB  TGT
0     A   1.0  NaN  1.0
1     A   1.0  NaN  NaN
2     B   NaN  1.0  NaN
3     B   1.0  NaN  NaN
4     B   1.0  NaN  NaN
5     C   NaN  NaN  NaN
6     C   1.0  NaN  1.0
7     C   1.0  NaN  NaN

需按以下规则设置TGT列的值:

  • 若每组中AAA=1.0的出现早于BBB=1.0,则为该组第一个出现AAA=1.0的行设置TGT=1;
  • 若每组中AAA=1.0的出现晚于BBB=1.0,则不设置任何行的TGT值;
  • 组内既无AAA=1.0也无BBB=1.0的行,保持TGT为NaN。

尝试了以下代码但未得到正确结果:

# Fill NaN values with a large negative value for comparison purposes
df.fillna(-9999, inplace=True)

# Filter rows where 'AAA' > 'BBB'
filtered_df = df.query('AAA > BBB')

# Group by 'group' column and get the first row of each group
filtered_df_group = filtered_df.groupby(df['group']).nth(0)
filtered_df_group ['TGT'] = 1

# set in the original df
df.loc[filtered_df_group.index, "TGT"] = filtered_df_group["TGT"]
解决方案

核心逻辑是先确定每组内AAA=1.0和BBB=1.0的首次出现顺序,再给符合条件的行赋值。以下是两种实现方式:

方式一:分步明确逻辑

import pandas as pd

# 初始化原始DataFrame
df = pd.DataFrame({
    'group': ['A', 'A', 'B', 'B', 'B', 'C', 'C', 'C'],
    'AAA': [1.0, 1.0, None, 1.0, 1.0, None, 1.0, 1.0],
    'BBB': [None, None, 1.0, None, None, None, None, None],
    'TGT': [1.0, None, None, None, None, None, 1.0, None]
})

# 1. 获取每组中AAA=1.0的首个索引,无则设为无穷大
first_aaa = df[df['AAA'] == 1.0].groupby('group').head(1).index
first_aaa_map = df.groupby('group').apply(
    lambda x: first_aaa[first_aaa.isin(x.index)].min() if not first_aaa.isin(x.index).empty else float('inf')
)

# 2. 获取每组中BBB=1.0的首个索引,无则设为无穷大
first_bbb = df[df['BBB'] == 1.0].groupby('group').head(1).index
first_bbb_map = df.groupby('group').apply(
    lambda x: first_bbb[first_bbb.isin(x.index)].min() if not first_bbb.isin(x.index).empty else float('inf')
)

# 3. 筛选出AAA先出现的组,找到这些组中第一个AAA=1.0的行
valid_groups = first_aaa_map[first_aaa_map < first_bbb_map].index
target_rows = df[(df['group'].isin(valid_groups)) & (df['AAA'] == 1.0)].groupby('group').head(1)

# 4. 设置TGT值,其余保持NaN
df['TGT'] = None
df.loc[target_rows.index, 'TGT'] = 1.0

方式二:简洁版(利用transform)

import pandas as pd

df = pd.DataFrame({
    'group': ['A', 'A', 'B', 'B', 'B', 'C', 'C', 'C'],
    'AAA': [1.0, 1.0, None, 1.0, 1.0, None, 1.0, 1.0],
    'BBB': [None, None, 1.0, None, None, None, None, None],
    'TGT': [1.0, None, None, None, None, None, 1.0, None]
})

# 标记每行是否是组内第一个AAA=1.0的行
df['is_first_aaa'] = df.groupby('group')['AAA'].transform(
    lambda x: x.eq(1.0).cumsum() == 1
) & df['AAA'].eq(1.0)

# 给每行标记组内第一个BBB=1.0的索引,无则设为极大值
df['first_bbb_idx'] = df.groupby('group')['BBB'].transform(
    lambda x: x[x.eq(1.0)].index.min() if x.eq(1.0).any() else float('inf')
)

# 根据条件设置TGT
df['TGT'] = df.apply(
    lambda row: 1.0 if row['is_first_aaa'] and row.name < row['first_bbb_idx'] else None,
    axis=1
)

# 清理临时列
df.drop(['is_first_aaa', 'first_bbb_idx'], axis=1, inplace=True)

最终输出

两种方式都会得到符合预期的结果:

group  AAA  BBB  TGT
0     A   1.0  NaN  1.0
1     A   1.0  NaN  NaN
2     B   NaN  1.0  NaN
3     B   1.0  NaN  NaN
4     B   1.0  NaN  NaN
5     C   NaN  NaN  NaN
6     C   1.0  NaN  1.0
7     C   1.0  NaN  NaN

内容的提问来源于stack exchange,提问作者Ashish kulkarni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 19:05:42