You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas将DataFrame中非NaN值替换为对应id列的值

Pandas:将DataFrame非NaN值替换为对应行id的高效实现

要实现将DataFrame中所有非NaN值替换为对应行的id列值,且避免低效的行列循环,我们可以利用Pandas的向量化操作完成,核心思路是通过广播id列结合条件替换函数实现。

实现代码

import pandas as pd
import numpy as np

# 原始DataFrame
df = pd.DataFrame({'id': ['X', 'Y', 'Z'],
                   'A': [1, np.nan, 0],
                   'B': [0, 0, np.nan],
                   'C': [np.nan, 1, 1]})

# 指定需要处理的列(排除id列)
target_cols = df.columns.drop('id')

# 广播id列并替换非NaN值
df[target_cols] = df[target_cols].where(
    df[target_cols].isna(), 
    df['id'].to_numpy()[:, None]
)

print(df)

代码解释

  1. 广播id列:df['id'].to_numpy()[:, None]将一维的id列转换为二维数组(形状为(行数, 1)),Pandas会自动将其广播到与目标列相同的形状,实现每行的id值对应到该行的所有列。
  2. 条件替换:df.where(condition, other)函数保留condition为True的元素,将condition为False的元素替换为other中的对应值。这里用df[target_cols].isna()作为条件,即保留原列中的NaN值,将非NaN值替换为对应行的id。

执行结果

输出的DataFrame完全符合需求:

id    A    B    C
0  X    X    X  NaN
1  Y  NaN    Y    Y
2  Z    Z  NaN    Z

替代方案(使用mask)

也可以用mask函数实现,逻辑与where相反(替换满足条件的元素):

df[target_cols] = df['id'].to_numpy()[:, None].mask(df[target_cols].isna())

这种向量化操作基于Pandas底层的C实现,相比Python原生循环效率提升显著,非常适合处理大数据集。

内容的提问来源于stack exchange,提问作者strohhut

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 22:21:35