You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Spark聚合函数时如何忽略指定字符(如N)

解决方法

可以通过先将col1中的字符"N"转换为null,再结合first函数的ignoreNulls=true参数来实现忽略"N"的效果,修改后的代码如下:

import org.apache.spark.sql.functions.{first, count, when}

df.groupBy("grp")
  .agg(
    first(when(df("col1") =!= "N", df("col1")), ignoreNulls = true),
    count("col2")
  ).show()

逻辑说明:

  • when(df("col1") =!= "N", df("col1"))会对col1列进行处理:当列值不等于"N"时保留原数值,等于"N"时返回null
  • 之后调用first函数并设置ignoreNulls=true,就会像忽略原生null值一样忽略掉被转换为null的"N"值,最终取到每组中第一个非"N"的col1值

内容的提问来源于stack exchange,提问作者John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 09:07:00