如何用Base R实现动态列分组统计、缺失值处理并替代pull函数
问题描述
现有如下数据集:
library(dplyr) set.seed(420) data <- data.frame(duration = c(3, 5, 6, 8, 10), rate = c(0.2, 0.5, 0.8, 0.85, 0.9)) %>% slice_sample(n = 705, replace = TRUE) %>% rowwise() %>% mutate(cured = sample(0:1, 1, prob = c(1 - rate, rate))) %>% select(-rate) head(data)
输出:
duration cured <dbl> <int> 1 10 1 2 10 1 3 5 1 4 3 1 5 10 0 6 10 1
原本使用dplyr按duration分组统计总数、cured列中1的个数及占比的代码如下:
data %>% group_by(duration) %>% summarise(n = n(), npos = sum(cured, na.rm = TRUE), rate = npos / n)
输出:
duration n npos rate <dbl> <int> <int> <dbl> 1 3 139 30 0.216 2 5 130 79 0.608 3 6 143 123 0.860 4 8 155 127 0.819 5 10 138 118 0.855
需求:
- 支持传入变量作为列名(而非固定
duration和cured) - 移除tidyverse依赖(用于包开发)
- 若
cured列存在缺失值,统计n和npos时完全忽略这些缺失值 - 找到Base R中替代
dplyr::pull(rate)获取列向量的方法
解决方案
1. 无依赖分组统计函数(满足需求1-3)
以下是完全基于Base R实现的分组统计函数,支持传入变量列名,自动过滤缺失值:
group_stats <- function(df, group_col, target_col) { # 过滤目标列存在缺失值的行(完全忽略这些样本) df_clean <- df[!is.na(df[[target_col]]), ] # 按分组列统计总样本数 n_total <- table(df_clean[[group_col]]) # 按分组列统计目标列中1的个数 n_positive <- tapply(df_clean[[target_col]], df_clean[[group_col]], sum, na.rm = TRUE) # 计算占比 pos_rate <- n_positive / n_total # 整理为结构化数据框并排序 result <- data.frame( group = as.integer(names(n_total)), n = as.integer(n_total), npos = as.integer(n_positive), rate = as.numeric(pos_rate) ) # 重命名分组列为传入的列名 colnames(result)[1] <- group_col # 按分组列排序 result <- result[order(result[[group_col]]), ] rownames(result) <- NULL return(result) }
2. 函数使用示例
# 构造带缺失值的测试数据(验证缺失值处理) set.seed(420) data_with_na <- data.frame(duration = c(3,5,6,8,10,3,NA,5), cured = c(1,0,1,1,0,NA,1,0)) # 调用函数,传入列名字符串变量 group_col <- "duration" target_col <- "cured" stats_result <- group_stats(data_with_na, group_col, target_col) print(stats_result)
输出:
duration n npos rate 1 3 1 1 1.0000000 2 5 2 0 0.0000000 3 6 1 1 1.0000000 4 8 1 1 1.0000000 5 10 1 0 0.0000000
3. Base R获取列向量(满足需求4)
直接使用Base R的索引方式即可替代dplyr::pull:
# 方式1:$符号索引(适合已知列名的情况) rate_vector1 <- stats_result$rate # 方式2:[[ ]]索引(适合列名存储在变量中的情况) col_name <- "rate" rate_vector2 <- stats_result[[col_name]]
内容的提问来源于stack exchange,提问作者jackahall
相关产品推荐
相关产品推荐

