在PySpark中复现SQL语句时遭遇列不可迭代问题求助
解决PySpark中substring参数类型不匹配问题
问题原因
你遇到的错误是因为PySpark的F.substring函数第三个参数在旧版本中仅支持字面量整数,而F.length()返回的是Column对象,无法直接作为参数传入。
解决方案
方案1:直接复用原SQL逻辑(最简便)
用F.expr()直接嵌入原SQL表达式,完全复用你之前的SQL逻辑,无需修改写法:
import pyspark.sql.functions as F df = df.withColumn( "location_id", F.when( df.location_type == "SUPPLIER", F.expr("SUBSTRING(location_id, 1, length(location_id)-3)") ).otherwise(df.location_id) )
方案2:针对特定格式优化(去除末尾-01)
如果你的供应商ID固定是XXXXXX-XX格式,也可以用substring_index直接按分隔符截取,避免长度计算:
import pyspark.sql.functions as F df = df.withColumn( "location_id", F.when( df.location_type == "SUPPLIER", F.substring_index(df.location_id, "-", 1) ).otherwise(df.location_id) )
方案3:正则表达式替换
用正则匹配并移除末尾3个字符,适合所有需要截断最后3位的场景:
import pyspark.sql.functions as F df = df.withColumn( "location_id", F.when( df.location_type == "SUPPLIER", F.regexp_replace(df.location_id, ".{3}$", "") ).otherwise(df.location_id) )
内容的提问来源于stack exchange,提问作者Chris
相关产品推荐
相关产品推荐

