You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言ggenealogy包中用列表批量获取个体五代祖先?

Batch Processing Ancestor Traversal with ggenealogy

Absolutely! You can batch-process your list of 2000 animals instead of running getAncestors() one by one — this is exactly where R's iterative functions shine. Here's how to do it smoothly:

Step 1: Prepare your list of animal IDs

First, make sure your 2000 animal IDs are stored in a vector or list. For example:

# Replace this with your actual list of 2000 animal IDs
animal_list <- c("5601T", "1234X", "5678Y", ...)

Step 2: Use lapply() to batch fetch ancestors

The lapply() function will iterate over every ID in your list, run getAncestors() for each, and return a list of results (one entry per animal). We'll also name the result list so you can easily map ancestors back to their original animal:

library(ggenealogy)
data(sbGeneal)

# Batch retrieve 5-generation ancestors for all animals
ancestors_results <- lapply(animal_list, function(animal_id) {
  getAncestors(animal_id, sbGeneal, 5)
})

# Name the result entries to match your animal IDs
names(ancestors_results) <- animal_list

Step 3 (Optional): Convert results to a data frame for easier analysis

If you prefer a tabular format (instead of a list) for downstream work (like filtering, counting, or visualization), you can convert the results to a data frame using dplyr:

library(dplyr)
library(tibble)

# Convert list to a tidy data frame
ancestors_df <- bind_rows(
  lapply(names(ancestors_results), function(animal_id) {
    tibble(
      animal_id = animal_id,
      ancestor_id = ancestors_results[[animal_id]]
    )
  })
)

Bonus: Speed up processing with parallel computing

Since you're dealing with 2000 animals, processing might take a bit of time. For a speed boost, use parallel processing. On Mac/Linux, you can use parallel::mclapply(); on Windows, foreach with a parallel backend works well:

Mac/Linux Parallel Example:

library(parallel)

# Use all but one core to avoid freezing your system
num_cores <- detectCores() - 1

ancestors_parallel <- mclapply(animal_list, function(animal_id) {
  getAncestors(animal_id, sbGeneal, 5)
}, mc.cores = num_cores)

names(ancestors_parallel) <- animal_list

Windows Parallel Example (using foreach):

library(foreach)
library(doParallel)

# Register parallel backend
cl <- makeCluster(detectCores() - 1)
registerDoParallel(cl)

# Run parallel processing
ancestors_parallel <- foreach(animal_id = animal_list, .packages = "ggenealogy") %dopar% {
  getAncestors(animal_id, sbGeneal, 5)
}

stopCluster(cl)
names(ancestors_parallel) <- animal_list

A quick note: If some animals don't have a full 5 generations of ancestors, getAncestors() will just return all available ancestors without throwing an error — so you don't need to handle missing cases separately.

内容的提问来源于stack exchange,提问作者ana_gg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 08:52:35