You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Scala计算DataFrame各列唯一值占比并打印可读结果

问题根因

你当前代码的问题是:println 操作是在Driver端构造Spark执行计划阶段执行的,此时countDistinct、count都是Spark的Column逻辑表达式,还没有触发行动算子执行实际计算,因此打印出来的是表达式的结构字符串,而非运行后的实际值。

优化后实现方案

该方案仅触发2次Spark作业(一次统计总记录数、一次聚合所有列的唯一值计数),性能远高于逐列统计的实现,代码如下:

import org.apache.spark.sql.functions._

def printUniqueColumnRatio(df: org.apache.spark.sql.DataFrame): Unit = {
  // 提前统计总记录数,避免重复计算
  val totalCount = df.count()
  // 构造所有列的去重计数聚合表达式
  val aggExprs = df.columns.map(colName => countDistinct(col(colName)).alias(s"${colName}_distinct"))
  // 触发聚合计算,一次性拿到所有列的唯一值计数
  val aggResult = df.agg(aggExprs.head, aggExprs.tail: _*).head()
  // 遍历格式化打印结果
  df.columns.foreach { colName =>
    val distinctCnt = aggResult.getAs[Long](s"${colName}_distinct")
    val ratio = distinctCnt.toDouble * 100 / totalCount
    // 格式化输出,列名左对齐占5位、计数左对齐占4位、百分比保留1位小数
    println(f"$colName%-5s $distinctCnt%-4d $ratio%.1f%%")
  }
}

如果需要保留返回结构化DataFrame的需求,可以用如下实现:

import org.apache.spark.sql.functions._
import org.apache.spark.sql.Row
import org.apache.spark.sql.types._

def uniqueColumnRatioToTotalNumberOfValues(df: org.apache.spark.sql.DataFrame): org.apache.spark.sql.DataFrame = {
  val totalCount = df.count()
  val aggExprs = df.columns.map(colName => countDistinct(col(colName)).alias(s"${colName}_distinct"))
  val aggResult = df.agg(aggExprs.head, aggExprs.tail: _*).head()
  // 构造结果行集合
  val resultRows = df.columns.map { colName =>
    val distinctCnt = aggResult.getAs[Long](s"${colName}_distinct")
    val ratio = distinctCnt.toDouble * 100 / totalCount
    Row(colName, distinctCnt, ratio)
  }.toSeq
  // 定义结果表结构
  val schema = StructType(Seq(
    StructField("column_name", StringType, nullable = false),
    StructField("distinct_count", LongType, nullable = false),
    StructField("ratio_percent", DoubleType, nullable = false)
  ))
  df.sparkSession.createDataFrame(resultRows, schema)
}

// 调用打印示例
uniqueColumnRatioToTotalNumberOfValues(df).collect().foreach { row =>
  println(f"${row.getAs[String]("column_name")}%-5s ${row.getAs[Long]("distinct_count")}%-4d ${row.getAs[Double]("ratio_percent")}%.1f%%")
}

内容的提问来源于stack exchange,提问作者joesan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 15:39:03