如何合并任务避免嵌套lapply?以R语言多参数绘图为例
lapply with Parameter Cartesian Products Great question! Nested lapply can get messy, especially when you’re looking to parallelize operations—overhead from nested parallel tasks can eat into performance. Let’s break down how to simplify this workflow and address your other concerns.
The core fix is to create a flat grid of all possible parameter combinations instead of nesting your lists. This lets you use a single mapping function to process every combination at once, which is perfect for parallelization since you’re working with a single, flat set of tasks.
Step 1: Generate All Parameter Combinations
For your example, we need every pairing of Species-split data and geom type. You can use tidyr::crossing (tidyverse-friendly) or base R’s expand.grid to build this grid:
Tidyverse Approach (Recommended for Readability)
library(tidyverse) # Split the data (we'll tackle the `drop` parameter behavior next) iris_ls <- split(iris, iris$Species, drop = TRUE) # Create a grid of all parameter pairs param_grid <- crossing( data = iris_ls, geom = c("bar", "box") ) # Use `pmap` to apply your function to every row of the grid plot_list <- param_grid %>% pmap(plot_fun)
Base R Approach
If you prefer sticking to base R, expand.grid works too—just make sure to handle the list column correctly:
iris_ls <- split(iris, iris$Species, drop = TRUE) geom_ls <- c("bar", "box") # Build the parameter grid param_grid <- expand.grid( data = iris_ls, geom = geom_ls, stringsAsFactors = FALSE ) # Use `mapply` with SIMPLIFY = FALSE to return a list of plots plot_list <- mapply(plot_fun, param_grid$data, param_grid$geom, SIMPLIFY = FALSE)
Scaling to Multiple Additional Parameters
This approach scales seamlessly if you have more parameters (like color schemes, theme tweaks, or other plot controls). Just add them to the parameter grid:
# Example with an extra `border_color` parameter param_grid <- crossing( data = iris_ls, geom = c("bar", "box"), border_color = c("steelblue", "darkorange") ) # Update your function call to use the new parameter plot_list <- param_grid %>% pmap(function(data, geom, border_color) { plot_fun(data, geom) + theme(panel.border = element_rect(color = border_color)) })
split(drop = TRUE) Behavior You’re not missing anything here—split(drop = TRUE) only removes empty factor levels from the resulting list (i.e., if a Species had no rows, it wouldn’t appear in iris_ls). It does not modify the factor levels inside each individual subset data frame.
To drop unused factor levels in each split subset, you’ll need to explicitly apply droplevels() to each data frame:
# Option 1: Tidyverse style iris_ls <- split(iris, iris$Species, drop = TRUE) %>% lapply(function(x) x %>% mutate(across(where(is.factor), droplevels))) # Option 2: Base R style iris_ls <- split(iris, iris$Species, drop = TRUE) %>% lapply(droplevels)
This is intentional behavior from split—it preserves the original factor structure unless you tell it otherwise, which is useful in many cases, but easy to adjust when you need to.
Instead of nesting parallel operations (which adds overhead), you now have a single flat list of tasks (each row in param_grid is one task). For parallelization, you can swap pmap with furrr::future_map (tidyverse parallel) or parallel::mclapply (base R parallel) directly on the grid, which is far more efficient than nested parallel calls.
内容的提问来源于stack exchange,提问作者Archymedes

