You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算DataFrame当前行与前后正值的间隔行数?

问题描述

给定如下DataFrame:

import pandas as pd
df = pd.DataFrame({"feature": [1, 0, 0, 0, 0, 1, 0, 1]})

原始数据输出:

feature
0        1
1        0
2        0
3        0
4        0
5        1
6        0
7        1

需要新增两列previous_feat和next_feat,分别记录当前行与上一个feature=1的行之间的间隔行数,以及与下一个feature=1的行之间的间隔行数。期望输出如下:

feature  previous_feat  next_feat
0        1            NaN        5.0
1        0            1.0        4.0
2        0            2.0        3.0
3        0            3.0        2.0
4        0            4.0        1.0
5        1            5.0        2.0
6        0            1.0        1.0
7        1            2.0        NaN

注:间隔可以是行数或索引差值,NaN可替换为0。

解决方案

核心思路是先定位所有feature=1的行索引,再通过索引计算每行与前后最近的1的间隔。使用numpy.searchsorted可以高效完成索引匹配:

import pandas as pd
import numpy as np

df = pd.DataFrame({"feature": [1, 0, 0, 0, 0, 1, 0, 1]})

# 获取所有feature=1的行索引
pos = df[df["feature"] == 1].index.values
idx = df.index.values

# 计算上一个1的间隔
prev_pos_idx = np.searchsorted(pos, idx, side="right") - 1
# 处理无前置1的情况(示例中仅首行)
prev_pos_idx[prev_pos_idx < 0] = np.nan
prev_pos = pos[prev_pos_idx]
df["previous_feat"] = idx - prev_pos

# 计算下一个1的间隔
next_pos_idx = np.searchsorted(pos, idx, side="left")
# 处理无后置1的情况(示例中仅末行)
next_pos_idx[next_pos_idx >= len(pos)] = np.nan
next_pos = pos[next_pos_idx]
df["next_feat"] = next_pos - idx

# 可选:将NaN替换为0
# df.fillna(0, inplace=True)

print(df)

代码解释

  1. 定位1的索引:pos存储所有feature=1的行索引,示例中为[0,5,7]。
  2. 上一个1的间隔计算:
    • np.searchsorted(pos, idx, side="right")找到每个索引在pos中右侧插入的位置,减1得到上一个1的索引位置。
    • 当前行索引减去上一个1的索引,得到间隔行数。
  3. 下一个1的间隔计算:
    • np.searchsorted(pos, idx, side="left")找到每个索引在pos中左侧插入的位置,得到下一个1的索引位置。
    • 下一个1的索引减去当前行索引,得到间隔行数。
  4. NaN处理:首尾的1没有前/后1,会生成NaN,可根据需求用fillna(0)替换。

内容的提问来源于stack exchange,提问作者imM4TT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 23:40:30