在R中分组计算时,如何跳过不满足函数条件的异常组?
解决分组计算相关系数时因数据不足报错的问题
针对你遇到的分组数据量不足导致cor.test报错终止代码的问题,可以用purrr包的possibly函数给cor.test添加容错机制,让它遇到不满足条件的分组时返回NA而非报错,从而让代码继续执行。
修改后的代码
library(dplyr) library(purrr) # 包装cor.test,出错时返回包含NA的列表 safe_cor_test <- possibly( .f = function(x, y) cor.test(x, y), # 出错时返回的默认结果,和cor.test的输出结构匹配 otherwise = list(estimate = NA, p.value = NA) ) t_h_korr <- t_h %>% group_by(Tag) %>% summarise(korr_test = list(safe_cor_test(Wassertemperatur, Lufttemperatur))) %>% mutate( rsquared = unlist(map(korr_test, "estimate")), pval = unlist(map(korr_test, "p.value")) )
关键说明
possibly函数会把原函数(这里是cor.test)转换成一个容错函数:当原函数正常运行时返回计算结果,出错时返回你指定的otherwise值。- 这里设置的
otherwise结构和cor.test的输出保持一致,这样后续用map提取estimate和p.value时不会出错,数据不足的分组会得到NA值。 - 如果需要更严格的过滤,也可以在
summarise之前先筛选出满足条件的分组,比如要求每组至少有2条非缺失数据:
t_h_korr <- t_h %>% group_by(Tag) %>% # 过滤掉非缺失数据少于2条的分组 filter(sum(!is.na(Wassertemperatur) & !is.na(Lufttemperatur)) >= 2) %>% summarise(korr_test = list(cor.test(Wassertemperatur, Lufttemperatur))) %>% mutate( rsquared = unlist(map(korr_test, "estimate")), pval = unlist(map(korr_test, "p.value")) )
这两种方法都能解决你的问题,第一种保留所有分组(不足的分组返回NA),第二种直接过滤掉不满足条件的分组,你可以根据需求选择。
内容的提问来源于stack exchange,提问作者nhydro
相关产品推荐
相关产品推荐

