You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark中搜索列内字符串并按子串规则有选择地替换为指定值

解决方案

你构造的sdf中String列是逗号分隔的字符串格式,我们可以通过Spark内置函数完成转换,全程无需自定义UDF,性能更优,仅要求Spark版本≥2.4,实现逻辑如下:

  1. 用split函数将逗号分隔的字符串拆分为数组
  2. 用transform函数遍历数组每个元素,按子串匹配规则做替换
  3. 用array_join将处理后的数组重新拼接为逗号分隔的字符串

完整实现代码

from pyspark.sql import functions as F
from pyspark.sql.functions import col, when, lit

# 定义匹配规则,后续新增规则只需修改此处即可
mapping = [
    ("EQU", "horse"),
    ("FEL", "cat"),
    ("BOS", "cow")
]

# 构建替换逻辑
replace_expr = F.transform(
    F.split(col("String"), ","),
    lambda x: when(x.contains(mapping[0][0]), lit(mapping[0][1]))\
             .when(x.contains(mapping[1][0]), lit(mapping[1][1]))\
             .when(x.contains(mapping[2][0]), lit(mapping[2][1]))\
             .otherwise(lit("other"))
)

# 执行转换
result_sdf = sdf.withColumn("String", F.array_join(replace_expr, ","))

# 输出验证
result_sdf.show(truncate=False)

输出结果

运行后输出和你预期完全一致:

+----+-----------------+
|Item|String           |
+----+-----------------+
|1   |horse,horse,other|
|2   |other,cat,other  |
|3   |cow,horse        |
+----+-----------------+

低版本Spark兼容方案

如果你使用的Spark版本低于2.4,可通过UDF实现相同逻辑:

from pyspark.sql.types import StringType

def replace_str(s):
    arr = s.split(",")
    res = []
    for item in arr:
        if "EQU" in item:
            res.append("horse")
        elif "FEL" in item:
            res.append("cat")
        elif "BOS" in item:
            res.append("cow")
        else:
            res.append("other")
    return ",".join(res)

replace_udf = F.udf(replace_str, StringType())
result_sdf = sdf.withColumn("String", replace_udf(col("String")))

内容的提问来源于stack exchange,提问作者user

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 18:06:07