You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas处理连续重复行保留首行索引并记录批次末尾索引的方法

实现方案

你可以通过识别连续重复批次边界分组聚合实现需求,通用写法适配任意列数的DataFrame:

完整代码

import pandas as pd

# 示例数据,可替换为你自己的DataFrame
df = pd.DataFrame(
    [[1, 1, 1], [2, 3, 4], [2, 3, 4], [1, 1, 1], [1, 1, 1], [1, 1, 1], [3, 3, 3]], 
    columns=["a", "b", "c"]
)

# 生成连续重复批次的组号:和上一行内容不同则组号+1
group_key = (df != df.shift()).any(axis=1).cumsum()

# 1. 取每个批次的首行,保留原始索引
first_rows = df.groupby(group_key).first()
# 2. 计算每个批次的末尾行索引
last_idx = df.reset_index().groupby(group_key)['index'].last().rename('last')
# 3. 拼接结果
df2 = pd.concat([last_idx, first_rows], axis=1).set_axis(first_rows.index)

结果验证

输出df2即可得到你需要的格式:

last  a  b  c
0     0  1  1  1
1     2  2  3  4
3     5  1  1  1
6     6  3  3  3

逻辑说明

  • df != df.shift() 逐元素对比当前行与上一行的差异,返回布尔矩阵
  • .any(axis=1) 识别新批次的起始行:当前行与上一行存在任意列不同时返回True
  • .cumsum() 对布尔值累加,同一连续批次的所有行都会分配到相同的组号
  • 分组后分别提取首行数据、批次末尾索引,拼接后即为目标结果

内容的提问来源于stack exchange,提问作者Matthew Strawbridge

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 18:39:00