求助:编写按组条件计数观测的错误提示函数
调查数据保密校验的R函数实现
核心思路
要满足需求,需要完成三个关键步骤:
- 按
county分组计算每个县的response总和 - 筛选出总和小于3的问题县
- 若存在问题县,生成包含具体县名的提示信息并触发错误/警告
完整实现代码(Base R版本)
check_confidentiality <- function(df) { # 按county分组统计response总和 county_response_sums <- aggregate(response ~ county, data = df, FUN = sum) # 筛选出未达保密阈值的县 problematic_counties <- county_response_sums$county[county_response_sums$response < 3] if (length(problematic_counties) > 0) { # 拼接错误提示信息 error_msg <- paste0("以下县未达保密阈值(response总和<3):", paste(problematic_counties, collapse = ", ")) # 触发错误(若需改为警告,将stop替换为warning即可) stop(error_msg) } else { message("所有县均符合保密要求,可正常共享数据。") } }
测试示例
用你提供的样本数据测试函数:
# 构造示例数据 df <- data.frame(county=c("A", "A", "A", "A", "A", "B", "B", "B", "B", "B", "C", "C", "C", "C", "C"), response=c(0, 1, 0, 1, 1, 0, 0, 1, 0, 1, 1, 1, 0, 1, 1), value=sample(20:100, 15, replace=TRUE)) # 调用校验函数 check_confidentiality(df)
运行后会抛出错误:Error in check_confidentiality(df) : 以下县未达保密阈值(response总和<3):B,完全匹配需求。
可选:dplyr版本(需加载dplyr包)
如果你习惯用tidyverse工具链,也可以用dplyr实现:
library(dplyr) check_confidentiality_dplyr <- function(df) { problematic_counties <- df %>% group_by(county) %>% summarise(total_response = sum(response), .groups = "drop") %>% filter(total_response < 3) %>% pull(county) if (length(problematic_counties) > 0) { error_msg <- paste0("以下县未达保密阈值(response总和<3):", paste(problematic_counties, collapse = ", ")) stop(error_msg) } else { message("所有县均符合保密要求,可正常共享数据。") } }
内容的提问来源于stack exchange,提问作者m-r-r-r
相关产品推荐
相关产品推荐

