You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取DataFrame每行首个非NaN值并将其余值替换为NaN

保留DataFrame每行首个非NaN值,其余设为NaN

示例输入

col1  col2  col3  col4
0   1.0   2.0   3.0   4.0
1   NaN   5.0   7.0  11.0
2   NaN   NaN  20.0   9.0
3   NaN   NaN   NaN  15.0

预期输出

col1  col2  col3  col4
0   1.0   NaN   NaN   NaN
1   NaN   5.0   NaN   NaN
2   NaN   NaN  20.0   NaN
3   NaN   NaN   NaN  15.0

实现方法

方法一:定位首个非NaN位置创建掩码

import pandas as pd
import numpy as np

# 构造示例DataFrame
df = pd.DataFrame({
    'col1': [1, np.nan, np.nan, np.nan],
    'col2': [2, 5, np.nan, np.nan],
    'col3': [3, 7, 20, np.nan],
    'col4': [4, 11, 9, 15]
})

# 获取每行第一个非NaN值的列索引
first_valid_cols = df.apply(pd.Series.first_valid_index, axis=1)

# 创建掩码矩阵,仅目标位置设为True
mask = pd.DataFrame(False, index=df.index, columns=df.columns)
for idx, col in first_valid_cols.items():
    mask.loc[idx, col] = True

# 应用掩码过滤,非目标位置替换为NaN
result = df.where(mask)

方法二:利用累积计数简化操作

# 更简洁的实现
result = df.mask(df.apply(lambda row: row.cumcount() != 0, axis=1))

原理:row.cumcount()会对每行的非NaN值从0开始递增计数,第一个非NaN值的计数为0,后续非NaN值计数≥1,通过mask方法将计数≠0的位置替换为NaN,刚好保留每行首个非NaN值。

内容的提问来源于stack exchange,提问作者angad2004

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 07:05:10