You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何筛选PySpark DataFrame中struct数组不含[Closed,Yes]键值对的行

PySpark 过滤实现方案

针对你给出的如下结构的DataFrame:

|-- name: string (nullable = true)
 |-- other_attr: array (nullable = true)
 |    |-- element: struct (containsNull = true)
 |    |    |-- key: string (nullable = true)
 |    |    |-- value: string (nullable = true)

你需要筛选出other_attr数组中不存在key=Closed且value=Yes的结构体、同时包含other_attr为null的行,可通过以下方法实现:


方法1:exists 高阶函数实现(Spark 2.4+ 推荐)

直接用高阶函数遍历数组判断是否存在匹配项,性能最优:

from pyspark.sql import functions as F

filtered_df = df.filter(
    # 保留other_attr为null的行,或者不存在匹配结构体的行
    F.col("other_attr").isNull() 
    | ~F.exists(
        F.col("other_attr"),
        lambda x: (x["key"] == "Closed") & (x["value"] == "Yes")
    )
)

方法2:SQL表达式写法(兼容性更好)

如果你使用的Spark版本不支持lambda形式的高阶函数调用,可以用expr直接写SQL表达式:

from pyspark.sql import functions as F

filtered_df = df.filter(
    F.expr("other_attr is null or not exists(other_attr, x -> x.key = 'Closed' and x.value = 'Yes')")
)

可选方法:filter + size 组合判断

也可以先过滤出数组中所有匹配的结构体,再通过数组长度判断是否存在匹配项:

from pyspark.sql import functions as F

filtered_df = df.filter(
    F.col("other_attr").isNull()
    | F.size(
        F.filter(F.col("other_attr"), lambda x: (x["key"] == "Closed") & (x["value"] == "Yes"))
    ) == 0
)

内容的提问来源于stack exchange,提问作者Reihan_amn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 08:24:04