You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas嵌套IF逻辑未获正确结果,请求排查问题

问题原因排查及修正

核心错误点

  • 函数判断逻辑失效:select函数里的if(travel and shop_high)是判断函数对象本身,而非函数执行后的结果。Python中函数对象属于“真值”,所以这个条件永远为True,完全没起到筛选Traveled==1且shopping=='High'的作用。
  • 全局DataFrame依赖错误:travel()和shop_high()直接操作全局的df,返回的是整列的布尔Series,而非针对当前行的判断结果,和apply逐行处理的逻辑完全不匹配。
  • 参数命名冲突:select函数的参数名是df,但apply(axis=1)传入的是单行Series,这个命名会和全局df混淆,逻辑上也不合理。
  • 边界情况未处理:未覆盖Travel_count ==3的场景(比如客户B),这类数据会返回None,但由于筛选条件失效,不符合要求的行(D、E、F)也会进入判断逻辑。

修正后的代码

方案1:简化逻辑(推荐)

直接用布尔索引先筛选符合条件的行,再生成结果:

import pandas as pd

df = pd.DataFrame({'customer':['A','B','C','D','E','F'],
                   'Traveled':[1,1,1,0,1,0],
                   'Travel_count':[2,3,5,0,1,0],
                   'country1':['UK','Italy','CA', '0','UK','0'],
                   'country2':['JP','IN','CO','0','EG','0'],
                   'shopping':['High','High','High','High','Medium','Medium']
                   })

# 筛选目标客户
filtered_df = df[(df['Traveled'] == 1) & (df['shopping'] == 'High')]

# 生成出行次数说明
def get_travel_note(row):
    if row['Travel_count'] > 3:
        return f"Customer {row['customer']} traveled more than 3 times"
    elif row['Travel_count'] < 3:
        return f"Customer {row['customer']} traveled less than 3 times"
    else:
        return f"Customer {row['customer']} traveled exactly 3 times"

filtered_df['note'] = filtered_df.apply(get_travel_note, axis=1)
result = filtered_df[['customer', 'shopping', 'note']]
print(result)

方案2:修复原函数结构

如果要保留原函数形式,修改如下:

import pandas as pd

df = pd.DataFrame({'customer':['A','B','C','D','E','F'],
                   'Traveled':[1,1,1,0,1,0],
                   'Travel_count':[2,3,5,0,1,0],
                   'country1':['UK','Italy','CA', '0','UK','0'],
                   'country2':['JP','IN','CO','0','EG','0'],
                   'shopping':['High','High','High','High','Medium','Medium']
                   })

# 修改为接收单行数据的判断函数
def travel(row):
    return row['Traveled'] == 1

def shop_high(row):
    return row['shopping'] == 'High'

def select(row):
    # 传入当前行执行判断
    if travel(row) and shop_high(row):
        if row['Travel_count'] > 3:
            return (row['customer'], row['shopping'], f"Customer {row['customer']} traveled more than 3 times")
        elif row['Travel_count'] < 3:
            return (row['customer'], row['shopping'], f"Customer {row['customer']} traveled less than 3 times")
        else:
            return (row['customer'], row['shopping'], f"Customer {row['customer']} traveled exactly 3 times")
    # 不符合条件返回None,后续dropna过滤
    else:
        return None

result = pd.DataFrame(list(df.apply(select, axis=1).dropna()))
result.columns = ['customer', 'shopping', 'note']
print(result)

预期输出

customer shopping                                       note
0        A      High        Customer A traveled less than 3 times
1        B      High       Customer B traveled exactly 3 times
2        C      High       Customer C traveled more than 3 times

(注:原预期提到仅包含A和C,若需排除B,可在筛选时添加& (df['Travel_count'] !=3))

内容的提问来源于stack exchange,提问作者MTALY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 22:00:58