如何使用tidyverse的map系列函数为数据框设置变量标签
解决方案
优化思路
- 先将数据字典转为命名查找向量,避免原循环中每次都要过滤字典的冗余操作,匹配效率从O(n)降到O(1)
- 用
purrr包的imap_dfr函数遍历数据集每一列,imap可以同时拿到列值和列名,非常适合当前场景
实现代码
library(tidyverse) # 原有基础数据 mydata <- tibble( a_1 = c(20,22, 13,14,44), a_2 = c(42, 13, 32, 31, 14), b = c("male", "female", "male", "female", "male"), c = c("Primary", "secondary", "Tertiary", "Primary", "Secondary") ) dictionary <- tibble( variable = c("a", "b", "c"), label = c("Age", "Gender", "Education"), type = c("mselect", "select", "select") ) # 构造标签查找向量 label_lookup <- deframe(dictionary[, c("variable", "label")]) # 用imap批量设置标签 mydata_with_label <- imap_dfr(mydata, ~{ # 提取列名前缀,匹配对应标签 attr(.x, "label") <- label_lookup[str_remove(.y, "_.*")] return(.x) })
效果验证
运行以下代码可以查看设置好的变量标签:
attr(mydata_with_label$a_1, "label") # 输出:"Age" attr(mydata_with_label$c, "label") # 输出:"Education"
如果变量数量很大,这个方案的速度会比原有循环快很多,核心优化点就是提前构造了查找表,避免了重复过滤字典的操作。
内容的提问来源于stack exchange,提问作者Moses
相关产品推荐
相关产品推荐

