Spark DataFrame基于布尔列使用when/otherwise语句报错排查
问题分析与修正方案
你的代码存在三个关键问题,导致运行报错:
- 逻辑运算符错误+优先级未明确:Spark中逻辑与/或需用
&&/||(而非位运算符&/|),且原逻辑未正确分组——应该先判断threshold_1或threshold_2为真,再和country的条件做与运算,必须加括号明确优先级。 - 布尔值判断冗余:布尔类型列无需写
== True,直接用列名即可表示“为真”的判断。 - 列名拼写错误:
threshold_2多了末尾空格,会导致Spark找不到对应列。
修正后的代码:
df = df.withColumn("Threshold_Filter", when((df["country"] == "INDIA") && (df["threshold_1"] || df["threshold_2"]), "Ind_country" ).otherwise("Dif_country"))
如果习惯用PySpark的col函数写法,可读性会更强:
from pyspark.sql.functions import col, when df = df.withColumn("Threshold_Filter", when((col("country") == "INDIA") & (col("threshold_1") | col("threshold_2")), "Ind_country") .otherwise("Dif_country"))
注:配合col函数时,Spark会将&/|解析为逻辑运算符,和直接用列索引的场景不同,但建议统一用&&/||避免混淆。
内容的提问来源于stack exchange,提问作者prasadrao thottempudi
相关产品推荐
相关产品推荐

