如何在Pandas DataFrame中按多条件映射生成新列
Pandas实现三条件映射创建新列
需求说明
根据col1的内容生成new_col,规则如下:
- 当
col1仅包含Ted时,new_col设为Ted - 当
col1同时包含Ted和Not Ted时,new_col设为Both - 当
col1仅包含Not Ted时,new_col设为Not Ted
输入输出示例
| col1 | new_col |
|---|---|
| Ted, Ted | Ted |
| Ted, Not Ted | Both |
| Not Ted, Not Ted | Not Ted |
| Not Ted, Ted | Both |
| Ted, Ted | Ted |
实现方法
方法1:嵌套np.where处理多条件
利用np.where的嵌套特性,逐层判断条件:
import pandas as pd import numpy as np df = pd.DataFrame({ 'col1': ['Ted, Ted', 'Ted, Not Ted', 'Not Ted, Not Ted', 'Not Ted, Ted', 'Ted, Ted'] }) df['new_col'] = np.where( df['col1'].str.contains('Ted') & df['col1'].str.contains('Not Ted'), 'Both', np.where( df['col1'].str.contains('Ted'), 'Ted', 'Not Ted' ) )
方法2:apply结合集合判断
将每行内容转为集合,通过集合元素直接匹配规则:
def map_col1(s): values = set(s.split(', ')) if {'Ted', 'Not Ted'}.issubset(values): return 'Both' elif 'Ted' in values: return 'Ted' else: return 'Not Ted' df['new_col'] = df['col1'].apply(map_col1)
方法3:矢量化字符串操作(高效推荐)
通过布尔条件组合直接赋值,避免循环开销:
has_ted = df['col1'].str.contains('Ted') has_not_ted = df['col1'].str.contains('Not Ted') df['new_col'] = 'Not Ted' df.loc[has_ted & ~has_not_ted, 'new_col'] = 'Ted' df.loc[has_ted & has_not_ted, 'new_col'] = 'Both'
内容的提问来源于stack exchange,提问作者Eisen
相关产品推荐
相关产品推荐

