Pandas数据分类结果异常:收入未达标却被标记为Y问题排查
问题原因与解决方案
你的代码出现不符合预期结果的核心原因是运算符优先级问题:Python中&(按位与)的优先级高于>、==这类比较运算符,导致条件表达式没有按照你预期的逻辑执行。
比如你写的row['total_income']>28000 & row['children']==0,会被Python解析为:row['total_income'] > (28000 & row['children'])
以你提到的第1行数据为例:
row['children']的值是1,28000 & 1的按位与结果为0row['total_income']为17932.802,显然大于0,因此这个条件判断结果为True,触发了test='Y'的赋值,完全偏离了你原本的逻辑。
解决方案
需要给每个比较表达式单独加括号,明确逻辑判断的顺序,你可以选择两种写法:
写法1:用&并添加括号
def assign_status(row): test='Negado' # Desempregado if row['income_type']=='unemployed': test='N' # Já devedor elif row['debt']==1: test='N' # Relação de renda e filhos. elif (row['total_income']>28000) & (row['children']==0): test='Y' # Teste recuperado elif (row['total_income']>31000) & (row['children']==1): test='Y' # Teste desconhecido elif (row['total_income']>34000) & (row['children']==2): test='Y' elif (row['total_income']>37000) & (row['children']==3): test='Y' elif (row['total_income']>40000) & (row['children']==4): test='Y' elif (row['total_income']>43000) & (row['children']==5): test='Y' else: test='N' return test
写法2:用and(更符合Python逻辑判断的常规写法)
def assign_status(row): test='Negado' # Desempregado if row['income_type']=='unemployed': test='N' # Já devedor elif row['debt']==1: test='N' # Relação de renda e filhos. elif row['total_income']>28000 and row['children']==0: test='Y' # Teste recuperado elif row['total_income']>31000 and row['children']==1: test='Y' # Teste desconhecido elif row['total_income']>34000 and row['children']==2: test='Y' elif row['total_income']>37000 and row['children']==3: test='Y' elif row['total_income']>40000 and row['children']==4: test='Y' elif row['total_income']>43000 and row['children']==5: test='Y' else: test='N' return test
两种写法都能解决你的问题,推荐使用and,它更直观,符合Python逻辑判断的常规用法。
内容的提问来源于stack exchange,提问作者Riuk2252
相关产品推荐
相关产品推荐

