如何将数据集通过管道传入pmap?管道传参错误排查
.传参的错误问题 我来帮你拆解这个问题的原因,以及给出更合适的解决方法:
错误原因:符号.的作用域冲突
你在管道里写pmap(list(country_code, crit_year), ~f_g_treat(dataset = ., country_code = ..1, crit_year = ..2 ))时,这里的.并不是管道前的dataset对象!
在purrr的匿名函数(~开头的写法)中,.代表的是当前迭代的整个输入组(或者说和..1等价,指第一个传入的参数),所以你实际上是把country_code的第一个数值(比如11)传给了f_g_treat的dataset参数,而filter()函数无法对数值型对象进行操作,自然就抛出了no applicable method for 'filter_' applied to an object of class "c('double', 'numeric')"的错误。
而你直接写dataset = dataset时,是明确引用了全局环境中的tibble对象,所以参数传递正确,函数能正常运行。
正确的写法方案
方案1:明确传递固定的数据集参数
因为dataset在pmap的迭代过程中是固定不变的,我们可以把它作为第三个固定参数传入pmap,避免和迭代参数混淆:
dataset <- dataset %>% mutate(treatment = pmap_chr( list(country_code, crit_year, list(dataset)), ~f_g_treat(dataset = ..3, country_code = ..1, crit_year = ..2) ))
这里用pmap_chr直接返回字符向量,省去了unlist()的步骤,更简洁。
方案2:用cur_data()获取当前管道中的数据集
在dplyr的管道语境下,cur_data()可以获取当前步骤的数据集,比直接写变量名更灵活(比如前面管道有修改数据集时也能正确获取):
dataset <- dataset %>% mutate(treatment = pmap_chr( list(country_code, crit_year), ~f_g_treat(dataset = cur_data(), country_code = ..1, crit_year = ..2) ))
方案3:改用更高效的匹配合并方式(推荐)
其实你的需求本质是给每个country匹配对应的crit_year,然后判断分组,完全不需要用pmap循环处理,用left_join合并匹配表的方式效率更高,也更符合tidyverse的风格:
# 先创建国家-临界年份的匹配表 country_crit_table <- tibble( country = country_code, crit_year = crit_year ) # 合并后直接计算treatment dataset <- dataset %>% left_join(country_crit_table, by = "country") %>% mutate(treatment = ifelse(yrbirth >= crit_year - 7, "Treat", "Contr")) %>% select(-crit_year) # 不需要保留crit_year的话可以删除这一行
这种方法避免了多次重复过滤整个数据集,数据量越大,效率优势越明显。
内容的提问来源于stack exchange,提问作者mugdi

