Pandas apply(lambda)嵌套条件判断生成新列全为0的原因求解
问题原因
代码运行后全返回0的核心问题是条件判断中的Python链式比较逻辑错误:
Python的连续比较语法会将x.date() in shortlongdates == True解析为等价的(x.date() in shortlongdates) and (shortlongdates == True)。由于shortlongdates是存储日期的列表,不可能和布尔值True相等,所以两个判断分支的条件永远为假,最终所有行都返回默认值0。
修正方案
方案1:直接修改原逻辑
去掉条件中冗余的== True判断即可,in运算本身就会返回布尔值,不需要额外比对:
data['ShortLongFlag'] = data['End DateTime'].apply(lambda x: -1 if (x.month == 3 and x.date() in shortlongdates) else (1 if (x.month == 10 and x.date() in shortlongdates) else 0))
方案2:更高效的向量化写法(推荐)
避免使用apply逐行遍历,用pandas向量化运算运行效率更高,尤其数据量大时差异明显:
import numpy as np # 提前构造判断条件 cond_mar = (data['End DateTime'].dt.month == 3) & (data['End DateTime'].dt.date.isin(shortlongdates)) cond_oct = (data['End DateTime'].dt.month == 10) & (data['End DateTime'].dt.date.isin(shortlongdates)) # 按条件赋值 data['ShortLongFlag'] = np.select([cond_mar, cond_oct], [-1, 1], default=0)
内容的提问来源于stack exchange,提问作者SlowlyLearning
相关产品推荐
相关产品推荐

