Pandas处理连续重复行保留首行索引并记录批次末尾索引的方法
实现方案
你可以通过识别连续重复批次边界分组聚合实现需求,通用写法适配任意列数的DataFrame:
完整代码
import pandas as pd # 示例数据,可替换为你自己的DataFrame df = pd.DataFrame( [[1, 1, 1], [2, 3, 4], [2, 3, 4], [1, 1, 1], [1, 1, 1], [1, 1, 1], [3, 3, 3]], columns=["a", "b", "c"] ) # 生成连续重复批次的组号:和上一行内容不同则组号+1 group_key = (df != df.shift()).any(axis=1).cumsum() # 1. 取每个批次的首行,保留原始索引 first_rows = df.groupby(group_key).first() # 2. 计算每个批次的末尾行索引 last_idx = df.reset_index().groupby(group_key)['index'].last().rename('last') # 3. 拼接结果 df2 = pd.concat([last_idx, first_rows], axis=1).set_axis(first_rows.index)
结果验证
输出df2即可得到你需要的格式:
last a b c 0 0 1 1 1 1 2 2 3 4 3 5 1 1 1 6 6 3 3 3
逻辑说明
df != df.shift()逐元素对比当前行与上一行的差异,返回布尔矩阵.any(axis=1)识别新批次的起始行:当前行与上一行存在任意列不同时返回True.cumsum()对布尔值累加,同一连续批次的所有行都会分配到相同的组号- 分组后分别提取首行数据、批次末尾索引,拼接后即为目标结果
内容的提问来源于stack exchange,提问作者Matthew Strawbridge
相关产品推荐
相关产品推荐

