You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark DataFrame中将字符串列转换为列表?

在PySpark DataFrame中将字符串列转换为列表的解决方案

针对你提到的把hashtags列从字符串转为列表的需求,根据字符串格式的不同,提供以下几种高效实现方式:

1. 处理JSON格式的字符串(推荐)

如果hashtags列的字符串是标准JSON数组格式(比如'["#python", "#spark"]'),直接用PySpark内置的from_json函数即可,无需自定义UDF,性能更优:

from pyspark.sql import SparkSession
from pyspark.sql.types import ArrayType, StringType
from pyspark.sql.functions import from_json

# 初始化SparkSession
spark = SparkSession.builder.appName("hashtag-convert").getOrCreate()

# 示例DataFrame
data = [("1", '["#python", "#spark"]'), ("2", '["#dataengineering"]')]
df = spark.createDataFrame(data, ["id", "hashtags"])

# 定义目标schema:字符串数组
array_schema = ArrayType(StringType())

# 转换列
df_converted = df.withColumn("hashtags_list", from_json(df["hashtags"], array_schema))

# 查看结果
df_converted.show(truncate=False)

如果你的字符串是单引号格式(比如"['#python', '#spark']"),需要先替换单引号为双引号,再用from_json:

from pyspark.sql.functions import regexp_replace

df_converted = df.withColumn("clean_hashtags", regexp_replace(df["hashtags"], "'", '"')) \
                 .withColumn("hashtags_list", from_json("clean_hashtags", array_schema))

2. 处理逗号分隔的简单字符串

如果hashtags列是用逗号直接分隔的字符串(比如"#python,#spark"),用split函数即可:

from pyspark.sql.functions import split

df_converted = df.withColumn("hashtags_list", split(df["hashtags"], ","))

3. 处理复杂自定义格式(不推荐,仅作为备选)

如果字符串格式非常特殊(比如不符合JSON规范),可以用udf结合Python的ast.literal_eval实现,注意这种方式性能不如内置函数:

from pyspark.sql.functions import udf
import ast

# 定义UDF
literal_eval_udf = udf(lambda x: ast.literal_eval(x), ArrayType(StringType()))

# 转换列
df_converted = df.withColumn("hashtags_list", literal_eval_udf(df["hashtags"]))

内容的提问来源于stack exchange,提问作者Mr.Bean

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 13:24:47