使用tidyverse的curly-curly语法时,如何处理字符串列名参数
解决方案
你遇到的问题是因为{{ }}(大括号插值)仅适用于裸列名(不带引号的列名),当传入字符串格式的列名时,sum("b")会被解释为对字符串"b"求和,自然会报类型错误。以下两种方法可以解决字符串列名作为参数的问题:
方法一:使用.data代词(推荐)
在tidyverse中,.data是指代当前数据框的内置代词,通过.data[[字符串列名]]可以直接引用对应列,完美适配字符串参数:
library(tidyverse) a = c(1,1,1,2,2) b = 1:5 c = 6:10 d = 9:13 dummy_data = tibble(a,b,c,d) calc_indicator = function(numerator, denominator){ dummy_data %>% group_by(a) %>% mutate( indicator_value = sum(.data[[numerator]]) / sum(.data[[denominator]]) ) } # 传入字符串参数测试 calc_indicator("b","d")
运行后会得到每组重复的指标值(因为是组内总和的比值,每组内所有行结果一致):
# A tibble: 5 × 5 # Groups: a [2] a b c d indicator_value <dbl> <int> <int> <int> <dbl> 1 1 1 6 9 0.273 2 1 2 7 10 0.273 3 1 3 8 11 0.273 4 2 4 9 12 0.316 5 2 5 10 13 0.316
方法二:使用rlang的符号转换
如果习惯用tidyeval语法,可以用sym()将字符串转换为列名符号,再用!!强制求值:
calc_indicator = function(numerator, denominator){ # 将字符串转为符号 num_sym <- sym(numerator) den_sym <- sym(denominator) dummy_data %>% group_by(a) %>% mutate( indicator_value = sum(!!num_sym) / sum(!!den_sym) ) } # 同样支持字符串参数 calc_indicator("b","d")
批量处理Excel中的多组指标
如果你的Excel中存储了多组分子分母的列名(比如一个包含numerator和denominator列的数据框),可以用purrr::map2批量生成结果:
# 模拟从Excel读入的指标定义 indicators <- tibble( indicator_name = c("ratio_b_d", "ratio_c_b"), numerator = c("b", "c"), denominator = c("d", "b") ) # 批量计算所有指标 result_list <- map2(indicators$numerator, indicators$denominator, calc_indicator) names(result_list) <- indicators$indicator_name # 查看其中一个结果 result_list$ratio_c_b
内容的提问来源于stack exchange,提问作者Tumaini Kilimba
相关产品推荐
相关产品推荐

