You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中按多条件映射生成新列

Pandas实现三条件映射创建新列

需求说明

根据col1的内容生成new_col,规则如下:

  • 当col1仅包含Ted时,new_col设为Ted
  • 当col1同时包含Ted和Not Ted时,new_col设为Both
  • 当col1仅包含Not Ted时,new_col设为Not Ted

输入输出示例

col1new_col
Ted, TedTed
Ted, Not TedBoth
Not Ted, Not TedNot Ted
Not Ted, TedBoth
Ted, TedTed

实现方法

方法1:嵌套np.where处理多条件

利用np.where的嵌套特性,逐层判断条件:

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'col1': ['Ted, Ted', 'Ted, Not Ted', 'Not Ted, Not Ted', 'Not Ted, Ted', 'Ted, Ted']
})

df['new_col'] = np.where(
    df['col1'].str.contains('Ted') & df['col1'].str.contains('Not Ted'),
    'Both',
    np.where(
        df['col1'].str.contains('Ted'),
        'Ted',
        'Not Ted'
    )
)

方法2:apply结合集合判断

将每行内容转为集合,通过集合元素直接匹配规则:

def map_col1(s):
    values = set(s.split(', '))
    if {'Ted', 'Not Ted'}.issubset(values):
        return 'Both'
    elif 'Ted' in values:
        return 'Ted'
    else:
        return 'Not Ted'

df['new_col'] = df['col1'].apply(map_col1)

方法3:矢量化字符串操作(高效推荐)

通过布尔条件组合直接赋值,避免循环开销:

has_ted = df['col1'].str.contains('Ted')
has_not_ted = df['col1'].str.contains('Not Ted')

df['new_col'] = 'Not Ted'
df.loc[has_ted & ~has_not_ted, 'new_col'] = 'Ted'
df.loc[has_ted & has_not_ted, 'new_col'] = 'Both'

内容的提问来源于stack exchange,提问作者Eisen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 20:55:04