如何在dplyr管道链中保存DataFrame的中间状态至新对象?
Great question—this is a common need when working with dplyr pipelines, and there are a couple clean ways to do it without breaking your workflow.
Method 1: Use the %T>% tee operator
The tee operator is made for exactly this scenario: it lets you perform a side effect (like saving your data to a new object) while passing the original data through to the next step in the pipe. Here's how you'd modify your code:
df <- read.csv(file) %>% mutate(....) %>% mutate(....) %>% # Save current state to df2 here %T>% { df2 <- . } %>% group_by(....) %>% arrange(var) %>% summary()
What's happening here? The %T>% takes the DataFrame from the previous mutate steps, runs the expression { df2 <- . } (where . refers to the current DataFrame), then returns the original DataFrame so the pipe can continue with group_by and beyond.
Method 2: Use curly braces with the regular %>%
If you prefer not to use the tee operator, you can use curly braces to wrap the assignment and explicitly return the DataFrame to keep the pipe going:
df <- read.csv(file) %>% mutate(....) %>% mutate(....) %>% { df2 <- . # Capture the current state in df2 . # Return the DataFrame to continue the pipe } %>% group_by(....) %>% arrange(var) %>% summary()
This works because expressions inside curly braces in a pipe return the last value—here, we return ., which is the DataFrame at that point in the pipeline.
Both methods work seamlessly with the tidyverse (since %T>% is part of magrittr, which is included with dplyr). Just pick whichever feels more readable to you!
内容的提问来源于stack exchange,提问作者Forge

