PySpark使用func.when()报错:TypeError: 'Column' object is not callable
问题根源:拼写错误引发的调用异常
你碰到的TypeError: 'Column' object is not callable错误,核心原因是把otherwise拼写成了otherwisw——最后一个字母应该是e而非w。
为什么会触发这个错误?
PySpark的Column对象有个特殊机制:当你访问它不存在的属性时,它会自动生成一个引用对应名称列的Column实例。比如你写.otherwisw,PySpark会误以为你要引用名为otherwisw的列,返回一个Column对象。而后续你尝试给这个对象传参数(.otherwisw(lit(''))),就相当于把Column对象当成函数来调用,这自然会触发"Column对象不可调用"的报错。
修正后的代码
只需要把拼写错误的otherwisw改成正确的otherwise即可:
import sys from pyspark.sql.window import Window from pyspark.sql import Row import pyspark.sql.functions as func from pyspark.sql import DataFrameStatFunctions as statFunc from pyspark.sql.functions import coalesce, current_date, current_timestamp, lit, unix_timestamp, from_unixtime, \ row_number, mean a_df = sqlContext.table('opssup_dev_wrk_ct.wrk_ct_ods_batch_derived_extnd_stg2') b_df = sqlContext.table('opssup_dev_wrk_ct.wrk_ct_sap_batch_specific_dates') fdsi_df = sqlContext.table('opssup_dev_wrk_ct.wrk_ct_sap_batch_specific_dates') ldsi_df = sqlContext.table('opssup_dev_wrk_ct.wrk_ct_sap_batch_specific_dates') fds_df = sqlContext.table('opssup_dev_wrk_ct.wrk_ct_sap_batch_specific_dates') temp22_df = a_df \ .join(b_df,(a_df.batch_number==b_df.Batch_Number)) \ .join(fdsi_df,(a_df.First_DSI_BATCH_NUMBER==fdsi_df.Batch_Number),"left_outer") \ .join(ldsi_df,(a_df.Last_DSI_BATCH_NUMBER==ldsi_df.Batch_Number),"left_outer") \ .join(fds_df,(a_df.First_DS_BATCH_NUMBER==fds_df.Batch_Number),"left_outer") \ .select( \ a_df.driving_batch_number, \ a_df.batch_number, \ b_df.Material_Group, \ a_df.mfg_stage_code, \ func.when(a_df.batch_number==a_df.First_DSI_BATCH_NUMBER,lit('0')) \ .otherwise(lit('')) # 此处修正了拼写错误 .alias('mfg_start_date') \ )
排查小技巧
以后再遇到这类"Column object is not callable"错误时,可以优先检查这几点:
- 是否不小心把
Column对象当成函数调用(比如多写了括号) - 是否拼写错了PySpark API的方法名(比如这次的
otherwise) - 是否误将列属性访问写成了函数调用
内容的提问来源于stack exchange,提问作者Sham
相关产品推荐
相关产品推荐

