You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Base R实现动态列分组统计、缺失值处理并替代pull函数

问题描述

现有如下数据集:

library(dplyr)
set.seed(420)
data <- data.frame(duration = c(3, 5, 6, 8, 10),
                   rate = c(0.2, 0.5, 0.8, 0.85, 0.9)) %>%
  slice_sample(n = 705, replace = TRUE) %>%
  rowwise() %>%
  mutate(cured = sample(0:1, 1, prob = c(1 - rate, rate))) %>%
  select(-rate)
head(data)

输出:

duration cured
     <dbl> <int>
1       10     1
2       10     1
3        5     1
4        3     1
5       10     0
6       10     1

原本使用dplyr按duration分组统计总数、cured列中1的个数及占比的代码如下:

data %>%
  group_by(duration) %>%
  summarise(n = n(),
            npos = sum(cured, na.rm = TRUE),
            rate = npos / n)

输出:

duration     n  npos  rate
     <dbl> <int> <int> <dbl>
1        3   139    30 0.216
2        5   130    79 0.608
3        6   143   123 0.860
4        8   155   127 0.819
5       10   138   118 0.855

需求:

  • 支持传入变量作为列名(而非固定duration和cured)
  • 移除tidyverse依赖(用于包开发)
  • 若cured列存在缺失值,统计n和npos时完全忽略这些缺失值
  • 找到Base R中替代dplyr::pull(rate)获取列向量的方法

解决方案

1. 无依赖分组统计函数(满足需求1-3)

以下是完全基于Base R实现的分组统计函数,支持传入变量列名,自动过滤缺失值:

group_stats <- function(df, group_col, target_col) {
  # 过滤目标列存在缺失值的行(完全忽略这些样本)
  df_clean <- df[!is.na(df[[target_col]]), ]
  
  # 按分组列统计总样本数
  n_total <- table(df_clean[[group_col]])
  # 按分组列统计目标列中1的个数
  n_positive <- tapply(df_clean[[target_col]], df_clean[[group_col]], sum, na.rm = TRUE)
  # 计算占比
  pos_rate <- n_positive / n_total
  
  # 整理为结构化数据框并排序
  result <- data.frame(
    group = as.integer(names(n_total)),
    n = as.integer(n_total),
    npos = as.integer(n_positive),
    rate = as.numeric(pos_rate)
  )
  # 重命名分组列为传入的列名
  colnames(result)[1] <- group_col
  # 按分组列排序
  result <- result[order(result[[group_col]]), ]
  rownames(result) <- NULL
  
  return(result)
}

2. 函数使用示例

# 构造带缺失值的测试数据(验证缺失值处理)
set.seed(420)
data_with_na <- data.frame(duration = c(3,5,6,8,10,3,NA,5),
                           cured = c(1,0,1,1,0,NA,1,0))

# 调用函数,传入列名字符串变量
group_col <- "duration"
target_col <- "cured"
stats_result <- group_stats(data_with_na, group_col, target_col)

print(stats_result)

输出:

duration n npos      rate
1        3 1    1 1.0000000
2        5 2    0 0.0000000
3        6 1    1 1.0000000
4        8 1    1 1.0000000
5       10 1    0 0.0000000

3. Base R获取列向量(满足需求4)

直接使用Base R的索引方式即可替代dplyr::pull:

# 方式1:$符号索引(适合已知列名的情况)
rate_vector1 <- stats_result$rate

# 方式2:[[ ]]索引(适合列名存储在变量中的情况)
col_name <- "rate"
rate_vector2 <- stats_result[[col_name]]

内容的提问来源于stack exchange,提问作者jackahall

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 18:00:43