You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断指定列表元素是否存在于Spark DataFrame的列表列中?

问题描述

给定如下Spark DataFrame模拟数据:

test_data = [('1', '[tech, fx]'),
             ('2', '[industry, computer]'),
             ('3', '[5G, Apple]')]
test_data = spark.sparkContext.parallelize(test_data).toDF(['id', 'text'])

需新增一列indicator,标记列表['fx', 'computer']中的任意单词是否存在于text列中——存在则标记为'1',不存在则标记为'0',期望输出如下:

result = [('1', '[tech, fx]', '1'),
          ('2', '[industry, computer]', '1'),
          ('3', '[5G, Apple]', '0')]
result = spark.sparkContext.parallelize(result).toDF(['id', 'text', 'indicator'])

解决方案

方法1:正则匹配(简单直接)

利用Spark的rlike函数结合正则表达式,匹配目标词汇的出现:

from pyspark.sql.functions import when, col

target_words = ['fx', 'computer']
# 构建正则:匹配任意目标词,加入单词边界避免误匹配长词汇
pattern = '|'.join([f'\\b{word}\\b' for word in target_words])

result_df = test_data.withColumn(
    'indicator',
    when(col('text').rlike(pattern), '1').otherwise('0')
)

# 查看结果
result_df.show()

方法2:解析为数组后检查(适合结构化文本)

如果text列的列表格式标准,可以先将其转换为数组,再检查是否包含目标词汇:

from pyspark.sql.functions import expr

result_df = test_data.withColumn(
    'indicator',
    expr("""
        CASE 
            WHEN EXISTS (
                SELECT 1 
                FROM UNNEST(split(trim(text, '[]'), ', ')) AS words 
                WHERE words IN ('fx', 'computer')
            ) 
            THEN '1' 
            ELSE '0' 
        END
    """)
)

# 查看结果
result_df.show()

该方法优势在于当目标词汇数量较多时,无需逐个编写判断条件,扩展性更强。

内容的提问来源于stack exchange,提问作者user19495470

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:01:11