You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame新增列问题:Desc3非空取其值,否则取Desc1值

问题与解决方案

问题描述

我有一个DataFrame,其中某列数据以分隔符\分隔,每行的分隔符数量为1个或2个,不固定。已通过以下代码将该列拆分为多列:

df[['Desc1','Desc2','Desc3']] = df['Descr'].str.split("\\", expand=True)

现在部分行的df['Desc3']有值,部分行的df['Desc1']有值。想要新增一列,当Desc3非空时取其值,否则取Desc1的值。尝试了以下代码但无法正常工作:当row['Desc3']非空时能正确赋值,但当row['Desc3']为空时,无法返回row['Desc1']的值。

def fn_name(row):
    if row['Desc3'] == '':
        return row['Desc1']
    else:
        return row['Desc3']

df['name'] = df.apply(fn_name,axis=1)

1. 修复现有代码的问题

问题根源是str.split拆分后,空列可能填充的是NaN而非空字符串'',原代码只判断了空字符串,漏掉了NaN的情况。可以修改判断逻辑:

方式一:同时判断NaN和空字符串

import pandas as pd

def fn_name(row):
    if pd.isna(row['Desc3']) or row['Desc3'] == '':
        return row['Desc1']
    else:
        return row['Desc3']

df['name'] = df.apply(fn_name, axis=1)

方式二:利用布尔判断简化逻辑

空字符串和NaN在布尔判断中都会被视为False,可以直接简写:

def fn_name(row):
    return row['Desc3'] if row['Desc3'] else row['Desc1']

df['name'] = df.apply(fn_name, axis=1)

2. 更高效的实现方式(替代apply)

apply是逐行遍历,在大数据集上性能较差,推荐使用pandas/numpy的内置矢量化函数:

方式一:使用np.where

import numpy as np

df['name'] = np.where(df['Desc3'].notna() & (df['Desc3'] != ''), df['Desc3'], df['Desc1'])

方式二:使用combine_first

该函数会优先取第一个序列的非空值,缺失时取第二个序列对应的值,需要先把空字符串替换为NaN:

df['name'] = df['Desc3'].replace('', np.nan).combine_first(df['Desc1'])

方式三:使用Series.where

df['name'] = df['Desc3'].where(df['Desc3'].notna() & (df['Desc3'] != ''), df['Desc1'])

内容的提问来源于stack exchange,提问作者Alhpa Delta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:40:14