Databricks中SparkR函数执行失败:count_distinct方法报错排查
问题与解决:SparkR中
count_distinct报错的处理 问题描述
非R语言使用者,因分析需求使用SparkR,运行含管道操作的代码时触发报错,提示无法为SparkDataFrame找到count_distinct的继承方法,尝试多种写法仍未解决,疑问是否由管道操作导致。
初始尝试代码
table %>% SparkR::select('station') %>% SparkR::count_distinct()
错误信息
Error in (function (classes, fdef, mtable) : unable to find an inherited method for function ‘count_distinct’ for signature ‘"SparkDataFrame"’
其他尝试代码
SparkR::select(table, table$STATION) %>% SparkR::countDistinct()
SparkR::select(table, table$STATION) %>% SparkR::count_distinct()
解决方法
管道操作%>%本身没有问题,报错根源是对SparkR中countDistinct/count_distinct的用法理解错误:
- SparkR中的
countDistinct是聚合函数,不能直接作为SparkDataFrame的方法调用,必须配合agg()方法使用,用来对指定列进行去重计数。
正确代码示例
写法一(使用字符串列名)
table %>% SparkR::select('station') %>% SparkR::agg(SparkR::countDistinct(SparkR::col('station')))
写法二(使用列对象引用)
SparkR::select(table, table$STATION) %>% SparkR::agg(SparkR::countDistinct(table$STATION))
补充说明
如果需要统计每个station分组后的记录数(非去重计数),可以改用groupBy+count的组合:
table %>% SparkR::select('station') %>% SparkR::groupBy('station') %>% SparkR::count()
内容的提问来源于stack exchange,提问作者ZachW
相关产品推荐
相关产品推荐

