如何修改eval_data函数与logic列表实现R批量dplyr数据处理?
Fixing the
eval_data Function and logic List for Batch Processing with pmap Let's walk through the issues with your current setup and fix them step by step to get your batch data processing working as intended:
1. Key Problems in the Original Code
Your current eval_data function and logic list have a few critical issues that break some scenarios:
- Directly using
eval(dplyr_logic)doesn't properly integrate dplyr expressions into the pipeline (you need quasiquotation tools fromrlanghere). - The
select({ select_vector })syntax doesn't safely unquote your column selection vector. - The
I()in yourlogiclist doesn't handle the "no transformation" case correctly for the pipeline. - Chaining
mutatecalls with%>%inside anexpr()causes evaluation errors.
2. Revised eval_data Function
Here's the corrected function, with explanations of each change:
eval_data <- function(data, dplyr_logic, select_vector) { data %>% # Use !! to unquote and evaluate the dplyr expression in the pipeline { if (!rlang::is_empty(dplyr_logic)) !!dplyr_logic else . } %>% # Use all_of() to safely select columns from the character vector dplyr::select(dplyr::all_of(select_vector)) }
What Changed:
- Quasiquotation with
!!: The bang-bang operator (!!) tells dplyr to treat yourdplyr_logicexpression as part of the pipeline, not a literal object. - Empty Logic Handling: The
if/elsecheck ensures that when there's no transformation needed, we just pass the original data through with.. - Safe Column Selection:
dplyr::all_of(select_vector)properly references column names from your vector, even for newly created columns likeNew_Column1.
3. Fixed logic List
We need to adjust how we define the expressions to play nicely with the revised function:
logic <- list( # Replace I() with an empty expression for the no-transformation case rlang::expr(), rlang::expr(mutate(New_Column1 = case_when(Sepal.Length > 7 ~ 'Big', Sepal.Length > 6 ~ 'Medium', TRUE ~ 'Small'))), # Combine multiple mutate steps into a single mutate() call instead of chaining %>% rlang::expr(mutate( New_Column2 = case_when(Sepal.Width > 3.5 ~ 'Big2', Sepal.Width > 3 ~ 'Medium2', TRUE ~ 'Small2'), New_Column3 = case_when(Petal.Width > 2 ~ 'Big3', Petal.Width > 1 ~ 'Medium3', TRUE ~ 'Small3') )), rlang::expr(filter(Sepal.Width > 3)) )
What Changed:
- Empty Expression for No Transformation:
rlang::expr()replacesI()becauseI()wraps the data object instead of acting as a no-op in the pipeline. - Combined Mutate Calls: Instead of chaining
mutatewith%>%insideexpr(), we put both transformations into onemutate()call. This ensures the expression evaluates correctly in the pipeline.
4. Run the Batch Processing
Now you can use pmap exactly as you intended to process all your data frames in bulk:
library(purrr) library(dplyr) library(rlang) # Your original data and select vector definitions (unchanged) dd <- list() dd$data <- list(mutate0 = iris, mutate1 = iris, mutate2= iris, filter1 = iris ) select_vec <- list( c("Species", "Sepal.Length"), c("Species", "New_Column1"), c("Species", "New_Column2", "New_Column3"), c("Species", "Sepal.Width") ) # Execute the batch processing results <- pmap(list(dd$data, logic, select_vec), eval_data) # Inspect the output lapply(results, head)
This will return a list of data frames, each transformed and filtered according to your specifications, with only the columns you selected retained.
内容的提问来源于stack exchange,提问作者MayaGans
相关产品推荐
相关产品推荐

