PySpark中使用instr操作字符串列时出现Column is not iterable错误
问题解决:PySpark中"Column is not iterable"错误修复
错误原因分析
出现"Column is not iterable"错误的核心问题是括号配对错误,导致when函数的语法结构被破坏:
- Approach 1中,
when的第一个分支内的substring语句未正确闭合括号,使得.otherwise被错误绑定到substring方法而非when函数,触发语法解析异常。 - 同时存在
substring参数逻辑错误:otherwise分支中截取最后6位时,第三个参数应为固定值6,而非整个字符串长度。 - Approach 2存在拼写错误(
expe_featr_sict_id应为expc_featr_sict_id)及全角符号(›需改为半角>)问题。
修正后的代码实现
方法1(修正版)
from pyspark.sql.functions import substring, length, upper, instr, when, col df.select( '*', when( instr(col('expc_featr_sict_id'), upper(col('sub_prod_underscored'))) > 0, substring( col('expc_featr_sict_id'), instr(col('expc_featr_sict_id'), upper(col('sub_prod_underscored'))) + length(col('sub_prod_underscored')) + 1, length(col('expc_featr_sict_id')) # 若需截取到字符串末尾,可省略第三个参数 ) ).otherwise( substring(col('expc_featr_sict_id'), length(col('expc_featr_sict_id')) - 6, 6) ).alias('design_form_factor_temp') ).show()
方法2(修正版)
import pyspark.sql.functions as f df2 = df.withColumn( 'design_form_factor_temp', f.when( f.instr(f.col('expc_featr_sict_id'), f.upper(f.col('sub_prod_underscored'))) > 6, f.substring( f.col('expc_featr_sict_id'), f.instr(f.col('expc_featr_sict_id'), f.upper(f.col('sub_prod_underscored'))) + f.length(f.col('sub_prod_underscored')) + 1, f.length(f.col('expc_featr_sict_id')) ) ).otherwise( f.substring(f.col('expc_featr_sict_id'), f.length(f.col('expc_featr_sict_id')) - 6, 6) ) )
关键修正点
- 确保
when(...).otherwise(...)的语法结构完整,所有括号正确配对。 substring截取最后6位时,第三个参数明确设为6,避免多余字符截取。- 修正拼写错误与全角符号问题,保证代码语法合规。
内容的提问来源于stack exchange,提问作者mr.data_engg
相关产品推荐
相关产品推荐

