You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在PySpark中复现SQL语句时遭遇列不可迭代问题求助

解决PySpark中substring参数类型不匹配问题

问题原因

你遇到的错误是因为PySpark的F.substring函数第三个参数在旧版本中仅支持字面量整数,而F.length()返回的是Column对象,无法直接作为参数传入。

解决方案

方案1:直接复用原SQL逻辑(最简便)

用F.expr()直接嵌入原SQL表达式,完全复用你之前的SQL逻辑,无需修改写法:

import pyspark.sql.functions as F

df = df.withColumn(
    "location_id",
    F.when(
        df.location_type == "SUPPLIER",
        F.expr("SUBSTRING(location_id, 1, length(location_id)-3)")
    ).otherwise(df.location_id)
)

方案2:针对特定格式优化(去除末尾-01)

如果你的供应商ID固定是XXXXXX-XX格式,也可以用substring_index直接按分隔符截取,避免长度计算:

import pyspark.sql.functions as F

df = df.withColumn(
    "location_id",
    F.when(
        df.location_type == "SUPPLIER",
        F.substring_index(df.location_id, "-", 1)
    ).otherwise(df.location_id)
)

方案3:正则表达式替换

用正则匹配并移除末尾3个字符,适合所有需要截断最后3位的场景:

import pyspark.sql.functions as F

df = df.withColumn(
    "location_id",
    F.when(
        df.location_type == "SUPPLIER",
        F.regexp_replace(df.location_id, ".{3}$", "")
    ).otherwise(df.location_id)
)

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 14:35:23