使用purrr::pmap()为多数据框设置列标签时遇错误求助
批量为带前缀数据框设置列标签的问题及解决
问题描述
有10多个结构相似的数据框,列名带有a_、b_等不同前缀,需要批量设置列标签。单独使用labelled::set_variable_labels()可以正常运行,但用purrr::pmap()批量处理时,报错提示some variables not found in x:y, z。
示例数据
df1 = tribble( ~a_age, ~a01edu, ~other_vars, 35, 17, 1, 41, 14, 2, 28, 12, 3, 68, 99, 4 ) df2 = tribble( ~b_age, ~b01edu, ~some_vars, 25, 10, 2, 52, 8, 1, 31, 20, 5 ) df3 = tribble( ~c_age, ~c01edu, 55, 16, 47, 11, 68, 16, 36, 6, 29, 16 )
可正常运行的单数据框设置代码
df1 = df1 |> labelled::set_variable_labels( .labels = list("a_age" = "Age", "a01edu" = "Highest education completed") )
批量处理代码及错误信息
批量处理代码:
df_list = list(df1, df2, df3) |> setNames(c("a", "b", "c")) params = tribble( ~x, ~y, ~z, "a", "a_age", "a01edu", "b", "b_age", "b01edu", "c", "c_age", "c01edu" ) pmap(params, function(x, y, z) { df_list[[x]] |> labelled::set_variable_labels( .labels = list(y = "Age", z = "Highest education completed") ) } )
错误信息:
<error/rlang_error> Error in `pmap()`: ℹ In index: 1. Caused by error in `var_label<-.data.frame`: ! some variables not found in x:y, z --- Backtrace: 1. purrr::pmap(...) 2. purrr:::pmap_("list", .l, .f, ..., .progress = .progress) 5. global .f(x = .l[[1L]][[i]], y = .l[[2L]][[i]], z = .l[[3L]][[i]], ...) 6. labelled::set_variable_labels(...) 8. labelled:::`var_label<-.data.frame`(`*tmp*`, value = .labels) 9. base::stop("some variables not found in x:", missing_names)
错误原因分析
核心问题是你在构建.labels列表时,直接使用y = "Age"会把y当作列名的字面量,而不是引用变量y中存储的字符串值。比如第一次迭代时,代码实际在找名为y和z的列,而不是a_age和a01edu,自然会提示变量不存在。
解决方案
方案1:修改pmap中的列表构建逻辑
使用rlang::set_names创建带动态名称的列表,让列表的键是y和z对应的字符串值:
library(purrr) library(labelled) library(tibble) df_list = list(df1, df2, df3) |> setNames(c("a", "b", "c")) params = tribble( ~x, ~y, ~z, "a", "a_age", "a01edu", "b", "b_age", "b01edu", "c", "c_age", "c01edu" ) # 修正后的pmap代码 result_list = pmap(params, function(x, y, z) { # 构建动态命名的标签列表 label_list = set_names( c("Age", "Highest education completed"), c(y, z) ) df_list[[x]] |> set_variable_labels(.labels = label_list) }) # 给结果列表命名(可选) result_list = setNames(result_list, c("a", "b", "c"))
方案2:更简洁的前缀匹配方式(无需手动写params)
既然数据框列名的前缀和数据框名称对应,且要设置标签的列有规律(*_age对应Age,*01edu对应教育程度),可以用imap遍历数据框列表,通过前缀匹配自动识别目标列:
result_list = imap(df_list, function(df, prefix) { # 匹配带前缀的age列和edu列 age_col = paste0(prefix, "_age") edu_col = paste0(prefix, "01edu") df |> set_variable_labels( .labels = list( !!age_col := "Age", !!edu_col := "Highest education completed" ) ) })
这里用!!(非标准求值)把字符串转为列名符号,实现动态赋值。
验证结果
可以通过var_label()查看设置后的标签:
var_label(result_list[["a"]]) # 输出: # $a_age # [1] "Age" # # $a01edu # [1] "Highest education completed" # # $other_vars # NULL
内容的提问来源于stack exchange,提问作者nightstand
相关产品推荐
相关产品推荐

