如何在PySpark中实现round(col(),col()),传递列作为舍入精度参数
问题解决方法
报错原因
低于3.0版本的PySpark中,pyspark.sql.functions.round的第二个参数仅支持整数字面量,不支持传入Column类型对象,因此传入列作为舍入精度参数时会触发column is not callable错误。
可行解决方案
- 方案1:使用
expr()包装SQL表达式(全版本兼容,优先推荐)
直接将原生SQL逻辑封装到expr函数中即可生效,代码示例:
from pyspark.sql.functions import expr from pyspark.sql.types import DecimalType # 对应需求的实现代码 df = df.withColumn( "rounded_cost_amt", expr("CAST(ROUND(CostAmt, COALESCE(CurrencyDecimalPlaceNum, 2)) AS decimal(23,6))") )
方案2:升级PySpark到3.0及以上版本
3.0及更高版本的PySpark已经优化了round函数的参数支持,允许第二个参数传入Column对象,升级后你原本的代码可以直接正常运行。方案3:自定义UDF实现(不推荐,性能劣于原生函数)
如果无法升级版本也不想使用expr,可以通过自定义UDF实现逻辑:
from pyspark.sql.functions import udf, col from pyspark.sql.types import DecimalType, IntegerType @udf(returnType=DecimalType(23,6)) def custom_round(value, precision): if value is None: return None use_precision = precision if precision is not None else 2 return round(value, use_precision) df = df.withColumn( "rounded_cost_amt", custom_round( col("CostAmt"), col("CurrencyDecimalPlaceNum").cast(IntegerType()) ) )
内容的提问来源于stack exchange,提问作者Priyanka Choudhari
相关产品推荐
相关产品推荐

