如何在R语言ggenealogy包中用列表批量获取个体五代祖先?
Absolutely! You can batch-process your list of 2000 animals instead of running getAncestors() one by one — this is exactly where R's iterative functions shine. Here's how to do it smoothly:
Step 1: Prepare your list of animal IDs
First, make sure your 2000 animal IDs are stored in a vector or list. For example:
# Replace this with your actual list of 2000 animal IDs animal_list <- c("5601T", "1234X", "5678Y", ...)
Step 2: Use lapply() to batch fetch ancestors
The lapply() function will iterate over every ID in your list, run getAncestors() for each, and return a list of results (one entry per animal). We'll also name the result list so you can easily map ancestors back to their original animal:
library(ggenealogy) data(sbGeneal) # Batch retrieve 5-generation ancestors for all animals ancestors_results <- lapply(animal_list, function(animal_id) { getAncestors(animal_id, sbGeneal, 5) }) # Name the result entries to match your animal IDs names(ancestors_results) <- animal_list
Step 3 (Optional): Convert results to a data frame for easier analysis
If you prefer a tabular format (instead of a list) for downstream work (like filtering, counting, or visualization), you can convert the results to a data frame using dplyr:
library(dplyr) library(tibble) # Convert list to a tidy data frame ancestors_df <- bind_rows( lapply(names(ancestors_results), function(animal_id) { tibble( animal_id = animal_id, ancestor_id = ancestors_results[[animal_id]] ) }) )
Bonus: Speed up processing with parallel computing
Since you're dealing with 2000 animals, processing might take a bit of time. For a speed boost, use parallel processing. On Mac/Linux, you can use parallel::mclapply(); on Windows, foreach with a parallel backend works well:
Mac/Linux Parallel Example:
library(parallel) # Use all but one core to avoid freezing your system num_cores <- detectCores() - 1 ancestors_parallel <- mclapply(animal_list, function(animal_id) { getAncestors(animal_id, sbGeneal, 5) }, mc.cores = num_cores) names(ancestors_parallel) <- animal_list
Windows Parallel Example (using foreach):
library(foreach) library(doParallel) # Register parallel backend cl <- makeCluster(detectCores() - 1) registerDoParallel(cl) # Run parallel processing ancestors_parallel <- foreach(animal_id = animal_list, .packages = "ggenealogy") %dopar% { getAncestors(animal_id, sbGeneal, 5) } stopCluster(cl) names(ancestors_parallel) <- animal_list
A quick note: If some animals don't have a full 5 generations of ancestors, getAncestors() will just return all available ancestors without throwing an error — so you don't need to handle missing cases separately.
内容的提问来源于stack exchange,提问作者ana_gg

