You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中groupBy统计NULL值遇无结果返回问题排查

代码中的明显错误及修正方案

这段PySpark代码存在关键逻辑错误,会导致计算结果异常甚至无有效输出:

  • 缺失值统计逻辑错误:
    你用count(when(col(c).isNull(), c))统计缺失值行数,但count()仅统计非null的表达式结果。当col(c)为null时,when返回的c本身也是null,count(null)会返回0,最终算出的缺失率全为0。正确的缺失值计数应该用sum(when(col(c).isNull(), 1).otherwise(0)),它会把每一个缺失值记为1并求和。

  • 需确认函数导入:
    如果没从pyspark.sql.functions导入col、count、when、lit这些函数,代码会直接报错,更无法返回结果。

修正后的代码:

from pyspark.sql.functions import col, count, when, lit, sum

columns = ['col1', 'col2', 'col3']

amount_missing_df = df.groupby('country').agg(
    *[(sum(when(col(c).isNull(), 1).otherwise(0))/count(lit(1))).alias(c) for c in columns]
)
amount_missing_df.show()

额外要检查:原始DataFrame df必须包含country列以及指定的col1/col2/col3,且存在非空数据。

内容的提问来源于stack exchange,提问作者Henri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 22:37:42