You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中提取每行首个非NaN值生成新列?

提取每行第一个非NaN值生成新列

原始数据

你的DataFrame如下:

|    |   a |   b |   c |
|---:|----:|----:|----:|
|  0 | nan | nan |   1 |
|  1 | nan |   2 | nan |
|  2 |   3 |   3 |   3 |

需求:新增一列d,取值为[1, 2, 3],且列的数量不固定(少于30列)。

你的现有尝试

你已经通过以下代码获取了每行第一个非NaN值所在的列名:

df.isna().apply(lambda x: x.idxmin(), axis=1)

输出结果:

0    c
1    b
2    a
dtype: object

基于现有结果的取值方法

可以通过遍历索引和列名的方式,提取对应位置的值:

col_names = df.isna().apply(lambda x: x.idxmin(), axis=1)
df['d'] = [df.loc[i, col] for i, col in enumerate(col_names)]

更简便的方法

Pandas提供了更直接的方式提取每行第一个非NaN值,无需先获取列名:

方法1:利用first_valid_index

通过first_valid_index定位每行第一个非NaN值的列索引,再提取对应值:

df['d'] = df.apply(lambda row: row[row.first_valid_index()], axis=1)

方法2:利用方向填充(bfill/ffill)

通过向左填充(bfill(axis=1))将每行第一个非NaN值填充到最左侧列,再提取第一列的值:

df['d'] = df.bfill(axis=1).iloc[:, 0]

或者通过向右填充(ffill(axis=1))后提取最后一列的值,效果一致:

df['d'] = df.ffill(axis=1).iloc[:, -1]

复现代码

import io
import pandas as pd 

df = pd.read_csv(io.StringIO(',a,b,c\n0,,,1\n1,,2,\n2,3,3,3\n'))

内容的提问来源于stack exchange,提问作者baxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 12:24:24