如何在R语言中实现可提前终止的并行版`Find`循环?
Great question! You’re exactly right that most parallel frameworks in R (like foreach or basic future workflows) operate in a "wait-for-all" mode—they don’t return until every task finishes. But we can absolutely build a parallel version of Find that grabs the first result meeting your criteria, and even cancels remaining tasks to save resources.
Here’s a practical implementation using the future and promises packages, which support asynchronous execution perfect for this use case:
Step 1: Set Up Dependencies
First, install and load the required packages:
install.packages(c("future", "promises")) library(future) library(promises)
Choose a parallel backend (adjust based on your OS; multisession works cross-platform, multicore is faster for Linux/macOS):
plan(multisession) # Switch to plan(multicore) if on Linux/macOS
Step 2: Build the Parallel Find Function
This function mimics the behavior of base R’s Find, but runs tasks in parallel and returns the first matching result immediately:
parallel_find <- function(X, FUN, cond = function(res) TRUE) { # Launch asynchronous futures for each input element futures <- lapply(X, function(x) future(FUN(x))) # Loop until we find a match or exhaust all tasks while (length(futures) > 0) { # Check which futures have finished completed <- which(sapply(futures, resolved)) if (length(completed) > 0) { # Iterate through completed tasks to check for matches for (i in completed) { result <- value(futures[[i]]) # If we found a match, cancel remaining tasks and return if (cond(result)) { lapply(futures[-i], cancel) # Optional but resource-efficient return(result) } # Remove non-matching completed tasks from the list futures <- futures[-i] } } # Short sleep to avoid hogging CPU Sys.sleep(0.05) } # Return NULL if no matches found return(NULL) }
Step 3: Test It Out
Let’s simulate slow, random tasks to see how it works. We’ll look for the first task that returns a number greater than 10:
# Simulate a slow task with random wait time and random output slow_random_task <- function(x) { Sys.sleep(runif(1, 0.5, 2)) # Wait 0.5-2 seconds sample(1:20, 1) # Return a random number between 1-20 } # Run the parallel find first_match <- parallel_find(1:10, slow_random_task, cond = function(res) res > 10) cat("First matching result:", first_match, "\n")
Key Notes:
- Canceling Tasks: The
cancel()call is optional, but it’s a good idea to include it—this stops any remaining incomplete tasks from wasting CPU/memory once you’ve got your answer. - Asynchronous Logic: The
resolved()function checks which tasks have finished, andvalue()retrieves their results without blocking (since we only call it on completed futures). - Flexibility: Adjust the
condargument to match your specific criteria—just like base R’sFind, it’s a predicate function that returnsTRUEwhen a result is a match.
内容的提问来源于stack exchange,提问作者user3603486

