如何将PySpark DataFrame的StringType列转换为ArrayType列
将Spark DataFrame的String类型数组列转为ArrayType
方法1:使用from_json函数(推荐)
你的列内容是标准JSON数组格式,直接用Spark内置的from_json函数即可完成类型转换,该方法严格遵循JSON规范,能处理复杂场景(比如元素包含特殊字符)。
Python 实现
from pyspark.sql.types import ArrayType, StringType from pyspark.sql.functions import from_json # 假设DataFrame名为df,待转换列名为str_array_col df = df.withColumn("array_col", from_json(df.str_array_col, ArrayType(StringType())))
Scala 实现
import org.apache.spark.sql.types.{ArrayType, StringType} import org.apache.spark.sql.functions.from_json // 假设DataFrame为df,待转换列名为str_array_col val df = df.withColumn("array_col", from_json($"str_array_col", ArrayType(StringType)))
方法2:正则分割(简易方案)
如果字符串格式固定且无复杂元素(比如元素不含", "分割符),可以通过正则去除首尾的[]后再分割:
Python 实现
from pyspark.sql.functions import regexp_replace, split df = df.withColumn("array_col", split(regexp_replace(df.str_array_col, r'^\[|\]$', ''), ', "'))
Scala 实现
import org.apache.spark.sql.functions.{regexp_replace, split} val df = df.withColumn("array_col", split(regexp_replace($"str_array_col", "^\\[|\\]$", ""), ", \"")
注意:方法2仅适用于简单场景,若数组元素包含逗号或特殊字符,会出现分割错误,优先使用方法1。
内容的提问来源于stack exchange,提问作者Chris_007
相关产品推荐
相关产品推荐

