You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于两列条件生成新列?Pandas语法错误求助

基于Pandas两列的条件逻辑生成新列

现有Pandas数据框df2,需基于col_1和col_2的条件逻辑生成名为Tag的新列。数据框定义如下:

import pandas as pd
df2 = pd.DataFrame({'NOTES': ["PREPAID_HOME_SCREEN_MAMO","SCREEN_MAMO",
                              "> Unable to connect internet>4G Compatible>Set",
                              "No>Not Barred>Active>No>Available>Others>",
                              "Internet Not Working>>>>Unable To Connect To"], 
     'col_1': ["voice", "voice","data","other","voice"],
     'col_2': ["DATA", "voice","VOICE","VOICE","voice"]})

尝试用以下代码实现逻辑,但出现语法错误:

df2['Tag'] =             
            if df['col_1']=='data':
                return "Yes"
            elif df['col_2']:
                return "Yes"
            else:
                return "No"

错误原因

Python原生的if-else语句无法直接处理Pandas的Series对象(即数据框的列),因为Series是批量数据集合,需要用Pandas/NumPy提供的批量处理方法实现条件逻辑。


正确实现方法

方法1:使用numpy.where(推荐,效率最优)

利用numpy.where实现向量化条件判断,适合简单逻辑,处理大规模数据时效率远高于循环或apply:

import numpy as np

# 条件:col_1等于'data' 或者 col_2非空(字符串非空即视为True)
df2['Tag'] = np.where(
    (df2['col_1'] == 'data') | (df2['col_2']),
    "Yes",
    "No"
)

方法2:使用apply函数(适合复杂多分支逻辑)

通过apply逐行处理数据,适合逻辑更复杂的场景,但处理大数据量时效率较低:

def get_tag(row):
    if row['col_1'] == 'data':
        return "Yes"
    elif row['col_2']:  # 判断col_2是否非空
        return "Yes"
    else:
        return "No"

df2['Tag'] = df2.apply(get_tag, axis=1)

方法3:使用Pandasloc索引赋值

先初始化新列为默认值,再通过索引定位满足条件的行并修改值:

# 先设置默认值为"No"
df2['Tag'] = "No"
# 满足条件的行设为"Yes"
df2.loc[(df2['col_1'] == 'data') | (df2['col_2']), 'Tag'] = "Yes"

内容的提问来源于stack exchange,提问作者Virendra Patel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 16:50:43