如何提取Pandas DataFrame中含NaN的行及其前后非NaN行?
提取含NaN的行及其上下相邻行的解决方案
核心思路
通过标记目标行的索引或生成布尔掩码,筛选出所有location列为NaN的行,以及它们的上、下相邻行,同时处理索引越界、连续NaN等边界情况。
方法一:基于索引的实现(直观易理解)
import pandas as pd import numpy as np # 示例数据 df = pd.DataFrame({'location': ['Seattle', np.nan, 'Portland', 'San Francisco'], 'time': ['2022-06-01 12:00:00', '2022-06-01 13:00:00', '2022-06-01 14:00:00', '2022-06-01 15:00:00']}) # 1. 定位所有NaN行的索引 nan_idx = df[df['location'].isna()].index # 2. 收集需要保留的索引:NaN行 + 上一行 + 下一行 keep_idx = set() for idx in nan_idx: keep_idx.add(idx) if idx > 0: keep_idx.add(idx - 1) if idx < len(df) - 1: keep_idx.add(idx + 1) # 3. 筛选并按原顺序排序 result = df.loc[sorted(keep_idx)] print(result)
方法二:布尔掩码实现(简洁高效)
利用Pandas的shift()生成移位后的NaN判断,直接生成筛选掩码:
import pandas as pd import numpy as np df = pd.DataFrame({'location': ['Seattle', np.nan, 'Portland', 'San Francisco'], 'time': ['2022-06-01 12:00:00', '2022-06-01 13:00:00', '2022-06-01 14:00:00', '2022-06-01 15:00:00']}) # 掩码规则:当前行是NaN / 上一行是NaN / 下一行是NaN mask = df['location'].isna() | df['location'].shift(1).isna() | df['location'].shift(-1).isna() result = df[mask] print(result)
边界场景处理说明
- 连续多个NaN:比如
[A, NaN, NaN, B],两种方法都会保留A、两个NaN、B,符合需求 - NaN位于首行:比如
[NaN, A, B],会保留NaN和A(自动跳过不存在的上一行) - NaN位于末行:比如
[A, B, NaN],会保留B和NaN(自动跳过不存在的下一行)
内容的提问来源于stack exchange,提问作者mabiel
相关产品推荐
相关产品推荐

