You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比Pandas分组后前组最后值与后组首个值并生成状态列

解决跨组对比生成status列的问题

我明白你的需求啦——你需要的不是对比同组内的首尾值,而是跨组对比:前一个id组的最后一个value,和当前id组的第一个value是否相同,然后给每一行标记对应的status。之前的代码方向不对,咱们换个思路来解决:

第一步:明确核心逻辑

要实现你的期望结果,我们需要:

  • 记录每个id组的第一个value和最后一个value
  • 对比当前组的第一个value与前一个组的最后一个value,得到当前组第一行的status
  • 每个组的最后一行,除了最后一个id组的最后一行标记na,其余都标记different
  • 第一个id组的第一行标记na

第二步:完整代码实现

import pandas as pd

# 初始化你的DataFrame
df = pd.DataFrame({'id' : [1,1,2,2,3,3,4,4,5,6,6,7,7,8,8,9,9],
                'value'  : ["GC", "GD", "GD", "GQ","GQ","GR","LA","LK","LK",
                           "HA","HE","HE","JB","JB","JF","JF","JJ"]})

# 1. 提取每个id组的首、尾value,同时给每个组分配顺序索引
group_summary = df.groupby('id').agg(
    first_value=('value', 'first'),
    last_value=('value', 'last')
).reset_index()
group_summary['group_order'] = group_summary.index  # 给组按出现顺序编号

# 2. 计算每个组和前一组的对比结果(用于组内第一行的status)
group_summary['prev_group_last'] = group_summary['last_value'].shift(1)
group_summary['first_row_status'] = group_summary.apply(
    lambda x: 'na' if pd.isna(x['prev_group_last']) 
              else 'same' if x['first_value'] == x['prev_group_last'] 
              else 'different',
    axis=1
)

# 3. 给原DataFrame添加组编号,方便匹配状态
df['group_order'] = df['id'].map(group_summary.set_index('id')['group_order'])

# 4. 给每一行生成对应的status
def assign_status(row):
    group_data = group_summary[group_summary['group_order'] == row['group_order']].iloc[0]
    # 获取当前组的行索引范围
    group_rows = df[df['group_order'] == row['group_order']].index
    # 第一行用跨组对比的结果
    if row.name == group_rows[0]:
        return group_data['first_row_status']
    # 最后一行:如果是最后一个组,标记na;否则标记different
    elif row.name == group_rows[-1]:
        return 'na' if row['group_order'] == group_summary['group_order'].max() else 'different'
    # 兼容组内有更多行的情况
    else:
        return 'different'

df['status'] = df.apply(assign_status, axis=1)

# 整理成期望的输出格式
final_result = df.drop('group_order', axis=1)[['id', 'value', 'status']]
print(final_result.to_string(index=False))

第三步:验证结果

运行代码后,输出的结果和你给出的期望完全一致:

id value   status
 1    GC       na
 1    GD different
 2    GD      same
 2    GQ different
 3    GQ      same
 3    GR different
 4    LA different
 4    LK different
 5    LK      same
 6    HA different
 6    HE different
 7    HE      same
 7    JB different
 8    JB      same
 8    JF different
 9    JF      same
 9    JJ       na

代码逻辑解释

  • 我们先通过groupby.agg提取每个组的首尾value,给组编号后就能用shift(1)轻松拿到前一组的最后一个value,实现跨组对比
  • 给原DataFrame的每一行匹配对应的组编号后,根据行在组内的位置(第一行/最后一行)分配对应的status
  • 最后整理掉临时列,得到符合要求的结果

内容的提问来源于stack exchange,提问作者deeni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 18:25:13