如何在NSE函数中通过冒号快捷语法获取...参数的列名
嗨,这个问题我熟!你遇到的核心问题是rlang::exprs(...)只会把用户输入的内容原样保存为表达式,像estimate:sharemoe这种tidyselect风格的连续列选择,本质是一个调用表达式,不是直接的列名向量,所以没法直接提取列名。解决的关键是用tidyselect工具包来解析这些选择表达式,拿到实际的列名和顺序。
下面是具体的解决方案,我会一步步拆解:
1. 理解问题根源
当你输入estimate:sharemoe时,rlang::exprs(...)得到的是list(estimate:sharemoe),这是一个call类型的对象,不是列名字符串。直接用它去设置因子水平肯定会出错,因为你需要的是实际存在的列名,而不是这个表达式。
2. 用tidyselect::eval_select()解析选择
tidyselect::eval_select()专门用来处理tidyverse风格的列选择语法,它能把estimate:sharemoe、starts_with("est")这类选择器转换成数据中对应的列名,还能严格保留用户输入的顺序。
3. 重构你的gather_arrange函数
修改后的函数会先解析选择的列,再执行gather操作,最后设置因子水平:
library(tidyr) library(rlang) library(tidyselect) gather_arrange <- function(data, ...) { # 解析用户输入的列选择,得到对应列的位置(带列名) selected_cols <- eval_select(expr(c(...)), data = data) # 提取列名,保持用户选择的顺序 col_names <- names(selected_cols) # 执行gather操作,用all_of()传入列名(兼容tidyselect语法) gathered_data <- gather(data, key = "type", value = "value", all_of(col_names)) # 设置type列的因子水平为选择的列名顺序 gathered_data$type <- factor(gathered_data$type, levels = col_names) return(gathered_data) }
如果你的tidyr版本比较新,更推荐用pivot_longer替代gather(因为gather已经处于维护状态),修改后的代码如下:
gather_arrange <- function(data, ...) { selected_cols <- eval_select(expr(c(...)), data = data) col_names <- names(selected_cols) gathered_data <- pivot_longer( data, cols = all_of(col_names), names_to = "type", values_to = "value" ) gathered_data$type <- factor(gathered_data$type, levels = col_names) return(gathered_data) }
4. 测试验证
用示例数据测试一下连续列选择:
# 构造测试数据 test_data <- tibble( region = c("North", "South", "East"), estimate = c(500, 600, 700), moe = c(25, 30, 35), sharemoe = c(0.05, 0.06, 0.07) ) # 使用连续列选择调用函数 result <- gather_arrange(test_data, estimate:sharemoe) # 查看type列的因子水平 levels(result$type) # 输出: [1] "estimate" "moe" "sharemoe"
这样就能完美保留你选择列的顺序,设置正确的因子水平啦!
内容的提问来源于stack exchange,提问作者camille
相关产品推荐
相关产品推荐

