You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取Pandas DataFrame中含NaN的行及其前后非NaN行?

提取含NaN的行及其上下相邻行的解决方案

核心思路

通过标记目标行的索引或生成布尔掩码,筛选出所有location列为NaN的行,以及它们的上、下相邻行,同时处理索引越界、连续NaN等边界情况。

方法一:基于索引的实现(直观易理解)

import pandas as pd
import numpy as np

# 示例数据
df = pd.DataFrame({'location': ['Seattle', np.nan, 'Portland', 'San Francisco'],
                   'time': ['2022-06-01 12:00:00', '2022-06-01 13:00:00', '2022-06-01 14:00:00', '2022-06-01 15:00:00']})

# 1. 定位所有NaN行的索引
nan_idx = df[df['location'].isna()].index

# 2. 收集需要保留的索引:NaN行 + 上一行 + 下一行
keep_idx = set()
for idx in nan_idx:
    keep_idx.add(idx)
    if idx > 0:
        keep_idx.add(idx - 1)
    if idx < len(df) - 1:
        keep_idx.add(idx + 1)

# 3. 筛选并按原顺序排序
result = df.loc[sorted(keep_idx)]
print(result)

方法二:布尔掩码实现(简洁高效)

利用Pandas的shift()生成移位后的NaN判断,直接生成筛选掩码:

import pandas as pd
import numpy as np

df = pd.DataFrame({'location': ['Seattle', np.nan, 'Portland', 'San Francisco'],
                   'time': ['2022-06-01 12:00:00', '2022-06-01 13:00:00', '2022-06-01 14:00:00', '2022-06-01 15:00:00']})

# 掩码规则:当前行是NaN / 上一行是NaN / 下一行是NaN
mask = df['location'].isna() | df['location'].shift(1).isna() | df['location'].shift(-1).isna()
result = df[mask]
print(result)

边界场景处理说明

  • 连续多个NaN:比如[A, NaN, NaN, B],两种方法都会保留A、两个NaN、B,符合需求
  • NaN位于首行:比如[NaN, A, B],会保留NaN和A(自动跳过不存在的上一行)
  • NaN位于末行:比如[A, B, NaN],会保留B和NaN(自动跳过不存在的下一行)

内容的提问来源于stack exchange,提问作者mabiel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 04:42:30