You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中如何将数组类型列转为小写并保持数组类型?

解决PySpark数组列转小写后保持数组类型的问题

问题原因

你自定义的UDF未指定返回类型,Spark默认将其推断为StringType,导致转换后的列变成字符串而非数组结构。

解决方案1:给UDF指定返回类型

显式声明UDF的返回类型为ArrayType(StringType()),让Spark明确识别返回值是字符串数组:

from pyspark.sql import functions as F
from pyspark.sql.types import ArrayType, StringType

def lower(token):
    return list(map(str.lower, token))

# 指定返回类型为字符串数组
lower_udf = F.udf(lower, returnType=ArrayType(StringType()))

df_mod1 = temp.withColumn('token', lower_udf("words"))

执行后token列会保持array<string>的类型,数据结构完全符合预期。

解决方案2:使用Spark内置函数(推荐)

Spark 2.4及以上版本支持transform函数,可直接对数组的每个元素应用转换逻辑,无需自定义UDF(内置函数性能远优于自定义UDF):

from pyspark.sql import functions as F

df_mod1 = temp.withColumn("token", F.transform("words", lambda x: F.lower(x)))

这种方式无需额外定义函数,通过内置的transform和lower函数直接完成转换,自动保持数组类型,代码更简洁高效。

验证结果

两种方法执行后,token列的schema均为:

|-- token: array (nullable = true)
 |    |-- element: string (containsNull = true)

数据输出与期望一致:

+---+--------------------------------------+
|id |token                                 |
+---+--------------------------------------+
|0  |[this, is, spark]                     |
|1  |[i, wish, java, could, use, case, classes]|
|2  |[data, science, is, cool]             |
|3  |[machine, learning]                   |
+---+--------------------------------------+

内容的提问来源于stack exchange,提问作者merkle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 20:57:23