You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PySpark DataFrame的StringType列转换为ArrayType列

将Spark DataFrame的String类型数组列转为ArrayType

方法1:使用from_json函数(推荐)

你的列内容是标准JSON数组格式,直接用Spark内置的from_json函数即可完成类型转换,该方法严格遵循JSON规范,能处理复杂场景(比如元素包含特殊字符)。

Python 实现

from pyspark.sql.types import ArrayType, StringType
from pyspark.sql.functions import from_json

# 假设DataFrame名为df,待转换列名为str_array_col
df = df.withColumn("array_col", from_json(df.str_array_col, ArrayType(StringType())))

Scala 实现

import org.apache.spark.sql.types.{ArrayType, StringType}
import org.apache.spark.sql.functions.from_json

// 假设DataFrame为df,待转换列名为str_array_col
val df = df.withColumn("array_col", from_json($"str_array_col", ArrayType(StringType)))

方法2:正则分割(简易方案)

如果字符串格式固定且无复杂元素(比如元素不含", "分割符),可以通过正则去除首尾的[]后再分割:

Python 实现

from pyspark.sql.functions import regexp_replace, split

df = df.withColumn("array_col", split(regexp_replace(df.str_array_col, r'^\[|\]$', ''), ', "'))

Scala 实现

import org.apache.spark.sql.functions.{regexp_replace, split}

val df = df.withColumn("array_col", split(regexp_replace($"str_array_col", "^\\[|\\]$", ""), ", \"")

注意:方法2仅适用于简单场景,若数组元素包含逗号或特殊字符,会出现分割错误,优先使用方法1。

内容的提问来源于stack exchange,提问作者Chris_007

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 05:04:54