You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark DataFrame行操作:拆分字符串列实现行展开

解决Spark DataFrame拆分逗号分隔列的问题

这是个很常见的需求,我们可以借助Spark的split和explode函数轻松实现,而且完全能保证新列的String类型,下面分Python和Scala两种常用场景给出方案:

Python 实现

首先导入需要的函数,然后分两步处理:拆分字符串为数组,再展开数组为多行:

from pyspark.sql.functions import split, explode, col

# 假设你的原始DataFrame叫df
# 第一步:将column_B按逗号拆分为String类型的数组
df_with_array = df.withColumn("column_B_array", split(col("column_B"), ","))

# 第二步:展开数组,同时重命名列得到目标结构
result_df = df_with_array.select(
    col("column_A").alias("column_new_A"),
    explode(col("column_B_array")).alias("column_new_B")
)

# 可以打印Schema确认类型
result_df.printSchema()

Scala 实现

逻辑和Python一致,只是语法略有不同:

import org.apache.spark.sql.functions.{split, explode, col}

// 假设原始DataFrame为df
val dfWithArray = df.withColumn("column_B_array", split(col("column_B"), ","))

val resultDF = dfWithArray.select(
  col("column_A").alias("column_new_A"),
  explode(col("column_B_array")).alias("column_new_B")
)

// 验证列类型
resultDF.printSchema()

额外说明

如果你的column_B里存在逗号后带空格的情况(比如"1, 12, 21"),可以把split的分隔符改成正则表达式",\\s*",这样会自动忽略逗号后的空格,避免拆分出带空格的字符串:

# Python 版本调整split参数
split(col("column_B"), ",\\s*")

这样处理后,column_new_A和column_new_B都会保持String类型,完全符合你的需求。

内容的提问来源于stack exchange,提问作者Dipanjan Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:18:34