PySpark中如何在contains函数返回True时返回列表中的匹配字符串?
在Spark中提取content列匹配的关键词
需求说明
需要新增match列,显示content字段中匹配到指定关键词列表的对应值,适配关键词嵌入字符串(如下划线连接、直接拼接)的场景,而非仅空格分隔的单词。
解决方案
方法1:提取第一个匹配的关键词(单匹配场景)
通过遍历关键词列表,依次检查是否存在于content中,取第一个匹配的关键词:
from pyspark.sql import functions as F data= [ (1,"john_trader@gmail.com"), (2, "lucas turism llc"), (3,"maryinvestor@gmail.com"), (4, "peter.anderson@gmail.com") ] df=spark.createDataFrame(data, ("id",'content')) words = ["trader", "turism", "investor"] # 构建contains列,判断是否存在匹配关键词 conditions = " or ".join([f"content like '%{word}%'" for word in words]) df2 = df.withColumn('contains', F.expr(conditions)) # 构建match列,提取第一个匹配的关键词 match_expr = F.coalesce(*[ F.when(F.col("content").like(f"%{word}%"), F.lit(word)) for word in words ]) df_final = df2.withColumn("match", match_expr) df_final.show()
方法2:提取所有匹配的关键词(多匹配场景)
如果存在多个关键词匹配的情况,可以收集所有匹配项并拼接:
from pyspark.sql import functions as F # 承接上述df2数据 # 构建匹配关键词数组并过滤出存在的项,再转为字符串 match_expr = F.array_join( F.filter( F.array(*[F.lit(word) for word in words]), lambda x: F.col("content").like(f"%{x}%") ), ", " ) # 将空字符串转为null df_final = df2.withColumn("match", F.expr("nullif(match, '')")) df_final.show()
代码说明
- 替换了原方案中按空格分割
content的逻辑,改为直接检查关键词是否为content的子串,适配下划线、直接拼接等各种包含场景。 - 方法1使用
coalesce取第一个匹配的关键词,适合单匹配场景;方法2通过filter收集所有匹配项,适合多匹配场景。
内容的提问来源于stack exchange,提问作者Ana Beatriz
相关产品推荐
相关产品推荐

