在作用于tibble的mutate()函数中可否访问完整tibble或分组tibble?
你可以直接用dplyr内置的cur_data()函数引用当前管道传入的数据集,不需要修改你现有myfunc的逻辑,调整后的完整代码如下:
library(dplyr) df1 = expand.grid(x1=1:2,x2=1:2,x3=1:2,x4=1:2,x5=1:2,x6=1:2) %>% mutate( x7 = sample(1:2,64,T), y1 = rnorm(64) ) df2 = expand.grid(x1=1:2,x2=1:2,x3=1:2,x4=1:2,x5=1:2,x6=1:2) %>% mutate( x7 = sample(1:2,64,T), y2 = rnorm(64) ) myfunc <- function(data){ data %>% mutate(key = paste(x1,x2,x3,x4,x5,x6)) %>% pull(key) } joined_df = df1 %>% mutate(y3 = runif(64)) %>% mutate(key = myfunc(cur_data())) %>% inner_join( df2 %>% mutate(y4 = runif(64)) %>% mutate(key = myfunc(cur_data())), by='key' )
cur_data()是dplyr 1.0.0版本后推出的工具,作用是在mutate、filter这类dplyr动词的求值环境中,获取当前正在操作的完整数据集,不会修改原数据的结构,刚好匹配你只需要返回key向量、不想通过函数重建整份数据框的需求。
如果你使用的是更早版本的dplyr,也可以直接用管道占位符.代替cur_data(),写法为mutate(key = myfunc(.)),效果完全一致。
内容的提问来源于stack exchange,提问作者Max Candocia
相关产品推荐
相关产品推荐

