使用dplyr按id和time统计每行负无穷值的平均占比
解决方法
要实现按id和time分组,计算每组内每行负无穷值的平均占比,核心思路是先计算每行的负无穷占比,再对分组内的行取平均值。以下是两种高效实现方式:
方法1:使用rowwise() + c_across()(直观易读)
library(tidyverse) dt <- data.frame(id = rep(1:3, each = 4), time = rep(1:4, times=3), x = c(1, 2, 1, -Inf, 2, -Inf,1, 1, 5, 1, 2, 1), y = c(2, -Inf, -Inf, -Inf, -Inf, 5, -Inf, 2, 1, 2, 2, 2)) dt %>% rowwise() %>% # 计算当前行中负无穷值的占比(排除id和time列) mutate(inf_prop = mean(c_across(-c(id, time)) == -Inf)) %>% group_by(id, time) %>% # 对分组内的行取平均占比 summarise(n = mean(inf_prop), .groups = "drop") %>% # 可选:过滤掉无负无穷的分组 filter(n > 0)
方法2:使用rowSums()(大数据集更高效)
如果你的实际数据有大量行,这种方法比rowwise()更快:
dt %>% # 计算每行负无穷的数量 / 目标列总数 mutate(inf_prop = rowSums(select(., -id, -time) == -Inf) / ncol(select(., -id, -time))) %>% group_by(id, time) %>% summarise(n = mean(inf_prop), .groups = "drop") %>% filter(n > 0)
结果说明
两种方法都会输出符合预期的结果:
id time n 1 1 2 0.5 2 1 3 0.5 3 1 4 1.0 4 2 1 0.5 5 2 2 0.5 6 2 3 0.5
原代码问题分析
你之前的尝试错误地在summarise()中使用across(),这会按列计算负无穷的占比,而不是按行计算后再对分组取平均。正确逻辑是先得到每行的负无穷占比,再对分组内的行做聚合。
内容的提问来源于stack exchange,提问作者Ph.D.Student
相关产品推荐
相关产品推荐

