You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark拆分分号分隔字符串生成数组并移除末尾空元素方法

PySpark处理带末尾分号的字符串拆分解决方案

方案1:先移除末尾分号再拆分(推荐)

仅针对末尾多余的分号做处理,不会影响字符串中间的正常内容,性能最优:

from pyspark.sql.functions import col, split, regexp_replace

data = data.withColumn(
    "newcolumn",
    # 先匹配替换字符串末尾的分号为空,再做拆分
    split(regexp_replace(col("column"), r";$", ""), ";")
)

方案2:拆分后过滤空元素

如果字符串中可能存在连续分号导致的多处空元素,可以拆分后统一移除所有空字符串:

from pyspark.sql.functions import col, split, array_remove

data = data.withColumn(
    "newcolumn",
    # 拆分后移除数组中所有空字符串元素
    array_remove(split(col("column"), ";"), "")
)

完整测试示例

可以直接运行以下代码验证效果:

from pyspark.sql import SparkSession
from pyspark.sql.functions import col, split, regexp_replace

# 初始化SparkSession
spark = SparkSession.builder.appName("split_test").getOrCreate()

# 构造测试数据
test_data = [
    ("511;520;611;",),
    ("322;620",),
    ("3;321;",),
    ("334;344",)
]
df = spark.createDataFrame(test_data, schema=["column"])

# 执行转换
df_result = df.withColumn(
    "newcolumn",
    split(regexp_replace(col("column"), r";$", ""), ";")
)

# 打印输出
df_result.show(truncate=False)

运行后输出结果和预期完全一致:

+------------+---------------+
|column      |newcolumn      |
+------------+---------------+
|511;520;611;|[511, 520, 611]|
|322;620     |[322, 620]     |
|3;321;      |[3, 321]       |
|334;344     |[334, 344]     |
+------------+---------------+

内容的提问来源于stack exchange,提问作者BADS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 14:45:03