You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于PySpark DataFrame其他列的值创建新字符串列?

PySpark DataFrame拼接列生成新字符串列的实现

在PySpark中不能直接用+运算符拼接字符串字面量和DataFrame列,需要使用Spark提供的字符串函数来实现,以下是几种可行方案:

方案1:使用format_string(语法最简洁)

format_string支持类似Python字符串格式化的语法,直接将列值插入指定位置:

from pyspark.sql.functions import format_string

# 生成新列New
df = df.withColumn("New", format_string("Hey there %s %s!", "Name", "Surname"))

方案2:使用concat + lit

通过lit()将普通字符串转换为列表达式,再用concat拼接所有元素:

from pyspark.sql.functions import concat, lit

df = df.withColumn(
    "New",
    concat(lit("Hey there "), "Name", lit(" "), "Surname", lit("!"))
)

方案3:使用concat_ws(适合多列按固定分隔符拼接场景)

先通过concat_ws用空格拼接Name和Surname,再和前后的固定字符串组合:

from pyspark.sql.functions import concat_ws, lit, concat

df = df.withColumn(
    "New",
    concat(lit("Hey there "), concat_ws(" ", "Name", "Surname"), lit("!"))
)

执行后输出结果

IdNameSurnameNew
1JohnJohnsonHey there John Johnson!
2AnnaMariaHey there Anna Maria!

内容的提问来源于stack exchange,提问作者Alcibiades

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 04:24:45