如何提取Pandas DataFrame每行非NaN值,全NaN行保留NaN
问题与解决方案
需求
提取Pandas DataFrame中每行的唯一非NaN值;若某一行所有元素均为NaN,则该行结果保留NaN。
原DataFrame
| a | b | c |
|---|---|---|
| NaN | NaN | ghi |
| NaN | def | NaN |
| NaN | NaN | NaN |
| abc | NaN | NaN |
| NaN | NaN | NaN |
期望输出
| Result |
|---|
| ghi |
| def |
| NaN |
| abc |
| NaN |
解决方案
方法1:高效向量化操作(适合大数据量)
利用stack()、groupby和reindex组合实现,避免逐行循环,性能更优:
import pandas as pd import numpy as np # 创建原DataFrame df = pd.DataFrame({ 'a': [np.nan, np.nan, np.nan, 'abc', np.nan], 'b': [np.nan, 'def', np.nan, np.nan, np.nan], 'c': ['ghi', np.nan, np.nan, np.nan, np.nan] }) # 生成结果 result_df = df.stack() \ .groupby(level=0).first() \ .reindex(df.index) \ .to_frame('Result') print(result_df)
步骤说明
df.stack():将DataFrame中的非NaN值转换为多层索引Series,原行号作为一级索引,原列名作为二级索引groupby(level=0).first():按原行号分组,提取每组的第一个值(即该行的唯一非NaN值)reindex(df.index):保留原DataFrame的所有行索引,全NaN的行会自动填充为NaNto_frame('Result'):将Series转换为指定列名的DataFrame
方法2:逐行处理(适合小数据量,逻辑直观)
用apply逐行遍历,判断并提取非NaN值:
result_df = df.apply( lambda row: row.dropna().iloc[0] if not row.isna().all() else np.nan, axis=1 ).to_frame('Result')
步骤说明
axis=1:指定按行处理row.isna().all():判断当前行是否全为NaNrow.dropna().iloc[0]:若存在非NaN值,提取第一个(也是唯一的)非NaN值
内容的提问来源于stack exchange,提问作者user14085914
相关产品推荐
相关产品推荐

