You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

是否存在纯Pandas写法等价于非Pandas实现的序列起止标记功能?

问题:纯Pandas写法实现连续非零非空值序列的起止标记

背景

在Stack Overflow回答「如何标记Pandas DataFrame某列中非空非零值序列的起止?」问题时,我提供了更通用的方案,但核心实现属于“非Pandas写法”;其他答案采用的是“纯Pandas写法”。

定义说明

  • 纯Pandas写法:无需自定义Python函数(包括lambda函数)即可实现需求;
  • 非Pandas写法:必须依赖自定义函数才能实现需求。

核心问题

尝试将基于itertools.groupby的非Pandas实现转为纯Pandas写法时,发现两者分组逻辑存在本质差异:

  • Python groupby:仅对连续的相同值分组,分散的相同值会被分为不同组;
  • Pandas groupby:会将所有相同值(无论是否连续)归为同一组,无法直接替代Python groupby的连续分组功能。

现需确认:是否存在纯Pandas写法,可实现与以下代码完全相同的功能?
功能要求:在DataFrame中,针对每个Cycle分组内的Value列,标记连续非零非空值序列的开始(标记为'A')、结束(标记为'B');若序列仅含单个值,则标记为'AB';其余位置标记为0。

原非Pandas实现代码

data = { 'Cycle': [1,1,1,1,1,2,2,2,2,2,3,3,3,3,3],
         'Value': [1,0,0,0,2,3,4,0,5,6,0,0,7,0,0]}  
df = pd.DataFrame(data)
from itertools import groupby
def getPOI(df):
    itrCV = zip(df.Cycle, df.Value)
    lstCV = list(zip(df.Cycle, df.Value)) # 仅用于测试
    lstPOI = []
    print('Python groupby结果:', [ ((c, v), list(g)) for (c, v), g in groupby(lstCV, lambda cv: 
                          (cv[0], cv[1]!=0 and not pd.isnull(cv[1]))) ]
         ) # 仅用于测试
    for (c, v), g in groupby(itrCV, lambda cv: 
                            (cv[0], not pd.isnull(cv[1]) and cv[1]!=0)):
        llg = sum(1 for item in g) # 避免创建列表
        if v is False: 
            lstPOI.extend([0]*llg)
        else: 
           lstPOI.extend(['A']+(llg-2)*[0]+['B'] if llg > 1 else ['AB'])
    return lstPOI
df["POI"] = getPOI(df)
print(df)
print('---')
print(df.POI.to_list())

代码输出结果

Cycle  Value POI
0       1      1  AB
1       1      0   0
2       1      0   0
3       1      0   0
4       1      2  AB
5       2      3   A
6       2      4   B
7       2      0   0
8       2      5   A
9       2      6   B
10      3      0   0
11      3      0   0
12      3      7  AB
13      3      0   0
14      3      0   0
---
['AB', 0, 0, 0, 'AB', 'A', 'B', 0, 'A', 'B', 0, 0, 'AB', 0, 0]

现有纯Pandas写法的缺陷

Scott Boston提供的以下纯Pandas写法,无法处理Cycle内分散的非零值序列:

mp = df.where(df!=0).groupby('Cycle')['Value'].agg([pd.Series.first_valid_index, 
                                            pd.Series.last_valid_index])
df.loc[mp['first_valid_index'], 'POI'] = 'A'
df.loc[mp['last_valid_index'], 'POI'] = 'B'
df['POI'] = df['POI'].fillna(0)

分组逻辑对比代码

df.Value = df.Value.where(df.Value!=0).where(pd.isnull, 1)
print(  'Pandas groupby结果:',
        df.groupby(['Cycle','Value'], sort=False).groups
) 

内容的提问来源于stack exchange,提问作者user7711283

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:01:39