You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas如何按指定条件拼接含NaN值的两个字符串列

实现代码

依赖导入与测试数据构造

import pandas as pd
import numpy as np

# 构造题目给出的示例输入数据
df = pd.DataFrame({
    'ColA': ['a', 'b', 'c', np.nan, np.nan],
    'ColB': ['a b c', 'a c', np.nan, 'd e f', np.nan]
})

核心逻辑实现

通过自定义行处理函数逐行判断符合规则的输出:

def calc_colc(row):
    val_a = row['ColA']
    val_b = row['ColB']
    # 两列均为空的情况
    if pd.isna(val_a) and pd.isna(val_b):
        return np.nan
    # 仅ColA为空,返回ColB
    if pd.isna(val_a):
        return val_b
    # 仅ColB为空,返回ColA
    if pd.isna(val_b):
        return val_a
    # 两列均非空时,判断ColA单词是否存在于ColB的单词列表中
    b_word_set = set(val_b.split())
    if val_a in b_word_set:
        return val_b
    # 不存在则拼接两列值
    return f"{val_a} {val_b}"

# 应用函数生成新列
df['ColC'] = df.apply(calc_colc, axis=1)

输出验证

执行print(df)即可得到和示例预期完全一致的结果:

ColA   ColB   ColC
0    a  a b c  a b c
1    b    a c  b a c
2    c    NaN      c
3  NaN  d e f  d e f
4  NaN    NaN    NaN

内容的提问来源于stack exchange,提问作者ScottC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 23:36:01