如何基于PySpark DataFrame其他列的值创建新字符串列?
PySpark DataFrame拼接列生成新字符串列的实现
在PySpark中不能直接用+运算符拼接字符串字面量和DataFrame列,需要使用Spark提供的字符串函数来实现,以下是几种可行方案:
方案1:使用format_string(语法最简洁)
format_string支持类似Python字符串格式化的语法,直接将列值插入指定位置:
from pyspark.sql.functions import format_string # 生成新列New df = df.withColumn("New", format_string("Hey there %s %s!", "Name", "Surname"))
方案2:使用concat + lit
通过lit()将普通字符串转换为列表达式,再用concat拼接所有元素:
from pyspark.sql.functions import concat, lit df = df.withColumn( "New", concat(lit("Hey there "), "Name", lit(" "), "Surname", lit("!")) )
方案3:使用concat_ws(适合多列按固定分隔符拼接场景)
先通过concat_ws用空格拼接Name和Surname,再和前后的固定字符串组合:
from pyspark.sql.functions import concat_ws, lit, concat df = df.withColumn( "New", concat(lit("Hey there "), concat_ws(" ", "Name", "Surname"), lit("!")) )
执行后输出结果
| Id | Name | Surname | New |
|---|---|---|---|
| 1 | John | Johnson | Hey there John Johnson! |
| 2 | Anna | Maria | Hey there Anna Maria! |
内容的提问来源于stack exchange,提问作者Alcibiades
相关产品推荐
相关产品推荐

