Win7下R3.4.1中foreach()调用system()的并行计算实现求助
Hey there, let's work through this problem together. The big hurdle here is that your loop relies on a shared starter file that gets modified every iteration—if you just naively throw this into parallel, you'll run into race conditions where multiple processes overwrite or read outdated versions of starter at the same time, leading to messed-up results.
Let's break this down into two key scenarios based on how your starter file is used, then walk through practical, Windows 7 + R 3.4.1-compatible solutions:
Scenario 1: Each bootstrap task uses the original starter (modifications don't need to pass between tasks)
This is the simpler case—if each run of your executable only needs the original starter file, and the modifications to starter are just side effects that don't affect other tasks, we can avoid conflicts by giving each parallel task its own copy of starter.
Step-by-step implementation:
Prepare your file lists
First, get a list of all 500 bootstrapped files (adjust the pattern to match your actual filenames):bootstrap_files <- list.files(path = getwd(), pattern = "boot_.*\\.csv", full.names = TRUE) stopifnot(length(bootstrap_files) == 500) # Verify we have all filesCreate unique
startercopies for each task
Make a separatestarterfile for each bootstrap task so they don't interfere with each other:original_starter <- "starter.txt" # Replace with your actual starter filename task_starters <- paste0("starter_task_", seq_along(bootstrap_files), ".txt") # Copy the original starter to each task-specific file invisible(file.copy(from = original_starter, to = task_starters))Set up a PSOCK cluster (Windows-compatible)
Windows doesn't support R'smclapply(forked processes), so we'll use a PSOCK cluster from theparallelpackage (built into R 3.4.1):library(parallel) # Use all but one CPU core to avoid slowing down your system num_cores <- detectCores() - 1 cl <- makePSOCKcluster(num_cores)Define your task function
Write a function that handles one bootstrap file and its pairedstartercopy. Make sure to add a unique prefix to the 3 output files to prevent overwrites:process_single_boot <- function(boot_path, starter_path) { # Extract a unique ID from the bootstrap filename for output prefixing task_id <- gsub(".*/boot_|\\.csv", "", boot_path) output_prefix <- paste0("output_", task_id) # Run your executable (replace this command with your actual program call) system( paste0( "your_program.exe --input ", boot_path, " --starter ", starter_path, " --output-prefix ", output_prefix ), wait = TRUE ) # Optional: Clean up the task-specific starter file if you don't need it file.remove(starter_path) }Run parallel tasks
Pair up your bootstrap files and starter copies, then execute them in parallel:# Create a list of paired arguments for each task task_args <- mapply(list, bootstrap_files, task_starters, SIMPLIFY = FALSE) # Run tasks across the cluster parLapply(cl, task_args, function(args) { process_single_boot(args[[1]], args[[2]]) })Clean up the cluster
Don't forget to shut down the cluster when you're done:stopCluster(cl)
Scenario 2: Each task depends on the modified starter from the previous task
If your loop is strictly sequential (Task 2 needs the starter modified by Task 1, Task 3 needs the starter modified by Task 2, etc.), full parallelization isn't possible. However, you can still speed things up with batch-based parallelism if you can group tasks into independent batches where the starter modifications only matter within a batch.
Example approach:
- Split your 500 bootstrap files into 10 batches of 50 files each.
- Process each batch sequentially:
- For Batch 1, start with the original
starter, run all 50 tasks in parallel using copies of the initialstarterfor the batch. - After Batch 1 finishes, update the master
starterfile with the final modification from the batch (you'll need to define how to aggregate modifications from the batch). - Repeat for each subsequent batch using the updated master
starter.
- For Batch 1, start with the original
This won't give you full 500-task parallelism, but it can still cut down runtime significantly compared to pure serial execution.
Key Notes for Windows 7 + R 3.4.1:
- Firewall permissions: Make sure R is allowed through Windows Firewall—PSOCK clusters need network access to communicate between processes.
- Absolute paths: Use full absolute paths for your executable, bootstrap files, and
startercopies to avoid "file not found" errors in parallel processes. - Output safety: Always add unique identifiers to your executable's output files to prevent parallel processes from overwriting each other's work.
内容的提问来源于stack exchange,提问作者Elizabeth

