You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中使用instr操作字符串列时出现Column is not iterable错误

问题解决:PySpark中"Column is not iterable"错误修复

错误原因分析

出现"Column is not iterable"错误的核心问题是括号配对错误,导致when函数的语法结构被破坏:

  • Approach 1中,when的第一个分支内的substring语句未正确闭合括号,使得.otherwise被错误绑定到substring方法而非when函数,触发语法解析异常。
  • 同时存在substring参数逻辑错误:otherwise分支中截取最后6位时,第三个参数应为固定值6,而非整个字符串长度。
  • Approach 2存在拼写错误(expe_featr_sict_id应为expc_featr_sict_id)及全角符号(›需改为半角>)问题。

修正后的代码实现

方法1(修正版)

from pyspark.sql.functions import substring, length, upper, instr, when, col

df.select(
    '*',
    when(
        instr(col('expc_featr_sict_id'), upper(col('sub_prod_underscored'))) > 0,
        substring(
            col('expc_featr_sict_id'),
            instr(col('expc_featr_sict_id'), upper(col('sub_prod_underscored'))) + length(col('sub_prod_underscored')) + 1,
            length(col('expc_featr_sict_id'))  # 若需截取到字符串末尾,可省略第三个参数
        )
    ).otherwise(
        substring(col('expc_featr_sict_id'), length(col('expc_featr_sict_id')) - 6, 6)
    ).alias('design_form_factor_temp')
).show()

方法2(修正版)

import pyspark.sql.functions as f

df2 = df.withColumn(
    'design_form_factor_temp',
    f.when(
        f.instr(f.col('expc_featr_sict_id'), f.upper(f.col('sub_prod_underscored'))) > 6,
        f.substring(
            f.col('expc_featr_sict_id'),
            f.instr(f.col('expc_featr_sict_id'), f.upper(f.col('sub_prod_underscored'))) + f.length(f.col('sub_prod_underscored')) + 1,
            f.length(f.col('expc_featr_sict_id'))
        )
    ).otherwise(
        f.substring(f.col('expc_featr_sict_id'), f.length(f.col('expc_featr_sict_id')) - 6, 6)
    )
)

关键修正点

  • 确保when(...).otherwise(...)的语法结构完整,所有括号正确配对。
  • substring截取最后6位时,第三个参数明确设为6,避免多余字符截取。
  • 修正拼写错误与全角符号问题,保证代码语法合规。

内容的提问来源于stack exchange,提问作者mr.data_engg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 04:22:39