You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中如何在contains函数返回True时返回列表中的匹配字符串?

在Spark中提取content列匹配的关键词

需求说明

需要新增match列,显示content字段中匹配到指定关键词列表的对应值,适配关键词嵌入字符串(如下划线连接、直接拼接)的场景,而非仅空格分隔的单词。

解决方案

方法1:提取第一个匹配的关键词(单匹配场景)

通过遍历关键词列表,依次检查是否存在于content中,取第一个匹配的关键词:

from pyspark.sql import functions as F

data= [
  (1,"john_trader@gmail.com"),
  (2, "lucas turism llc"),
  (3,"maryinvestor@gmail.com"),
  (4, "peter.anderson@gmail.com")
]
df=spark.createDataFrame(data, ("id",'content'))

words = ["trader", "turism", "investor"]

# 构建contains列,判断是否存在匹配关键词
conditions = " or ".join([f"content like '%{word}%'" for word in words])
df2 = df.withColumn('contains', F.expr(conditions))

# 构建match列,提取第一个匹配的关键词
match_expr = F.coalesce(*[
    F.when(F.col("content").like(f"%{word}%"), F.lit(word))
    for word in words
])

df_final = df2.withColumn("match", match_expr)
df_final.show()

方法2:提取所有匹配的关键词(多匹配场景)

如果存在多个关键词匹配的情况,可以收集所有匹配项并拼接:

from pyspark.sql import functions as F

# 承接上述df2数据
# 构建匹配关键词数组并过滤出存在的项,再转为字符串
match_expr = F.array_join(
    F.filter(
        F.array(*[F.lit(word) for word in words]),
        lambda x: F.col("content").like(f"%{x}%")
    ),
    ", "
)

# 将空字符串转为null
df_final = df2.withColumn("match", F.expr("nullif(match, '')"))
df_final.show()

代码说明

  • 替换了原方案中按空格分割content的逻辑,改为直接检查关键词是否为content的子串,适配下划线、直接拼接等各种包含场景。
  • 方法1使用coalesce取第一个匹配的关键词,适合单匹配场景;方法2通过filter收集所有匹配项,适合多匹配场景。

内容的提问来源于stack exchange,提问作者Ana Beatriz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 03:40:22