使用datar包(dplyr)统计DataFrame非空值遇报错问题
解决datar分组统计非空值时的
count报错问题 问题背景
用datar处理DataFrame时,已经通过group_by+summarise实现了两个分组统计项:
distinct:b列不同值的数量(用n_distinct(f.b))total_by_group:分组的总行数(用n())
但尝试用count(f.b)统计b列非空值数量时,触发了NotImplementedByCurrentBackendError错误,需要得到包含非空值计数的最终结果。
现有可运行代码及输出
from datar.all import * df = tibble(a=c("a","b","a","b","a"),b=c(1,1,None,1,2)) # 当前正常执行的分组统计 current_result = df >> group_by(f.a) >> summarise(distinct=n_distinct(f.b) ,total_by_group=n()) print(current_result)
输出:
a distinct total_by_group <object> <int64> <int64> 0 a 2 3 1 b 1 2
报错代码及信息
尝试添加非空值统计的代码:
df >> group_by(f.a) >> summarise(count(f.b))
触发的错误:
NotImplementedByCurrentBackendError: 'count' (data type: SeriesGroupBy, backends: numpy, pandas)
解决方案
datar的count函数目前不支持对分组后的Series直接调用,我们可以改用sum(~is.na(f.b))来统计非空值数量——通过is.na(f.b)标记空值,取反后得到非空值的布尔序列,求和即可得到非空值的个数。
修正后的代码及期望输出
from datar.all import * df = tibble(a=c("a","b","a","b","a"),b=c(1,1,None,1,2)) # 添加非空值统计的正确代码 final_result = df >> group_by(f.a) >> summarise( distinct=n_distinct(f.b), total_by_group=n(), count=sum(~is.na(f.b)) ) print(final_result)
输出:
a distinct total_by_group count <object> <int64> <int64> <int64> 0 a 2 3 2 1 b 1 2 2
内容的提问来源于stack exchange,提问作者Guillaume Lombard
相关产品推荐
相关产品推荐

