如何从sampleManifest提取配对样本Role并合并为新列?
Got it, let's tackle this problem step by step. You're already thinking in the right direction with left_join—we just need to expand that approach to pull in the role for both samples in each pair, then combine those roles into your desired Roles column.
Step 1: Prep Your Role Lookup
First, let's create a simplified version of your sampleManifest that only includes the columns we need: SampleName (to match against ID1/ID2) and Role. This keeps our joins clean and avoids extra columns cluttering up the process.
library(tidyverse) # Load dplyr and stringr for this workflow # Create a lookup table for sample roles role_lookup <- sampleManifest %>% select(SampleName, Role)
Step 2: Join Roles for Both Samples
Next, we'll join this lookup table to kinshipEstimates twice: once to get the role for ID1, and again to get the role for ID2. We'll rename these columns to Role1 and Role2 so we can tell them apart.
kinship_with_roles <- kinshipEstimates %>% # Join to get Role for ID1 left_join(role_lookup, by = c("ID1" = "SampleName")) %>% rename(Role1 = Role) %>% # Join again to get Role for ID2 left_join(role_lookup, by = c("ID2" = "SampleName")) %>% rename(Role2 = Role)
Step 3: Concatenate Roles into a Single Column
Now we'll use str_c() (from stringr) to combine Role1 and Role2 with a hyphen separator, then rearrange the columns to match your desired output format.
kinship_with_roles <- kinship_with_roles %>% mutate(Roles = str_c(Role1, Role2, sep = "-")) %>% # Reorder columns to match your target structure select(FID, ID1, ID2, Roles, Kinship, Relationship)
Combined One-Liner (Optional)
If you prefer a more concise workflow, you can chain all these steps together without creating a separate lookup table:
kinship_with_roles <- kinshipEstimates %>% left_join(sampleManifest %>% select(SampleName, Role), by = c("ID1" = "SampleName")) %>% rename(Role1 = Role) %>% left_join(sampleManifest %>% select(SampleName, Role), by = c("ID2" = "SampleName")) %>% rename(Role2 = Role) %>% mutate(Roles = str_c(Role1, Role2, sep = "-")) %>% select(FID, ID1, ID2, Roles, Kinship, Relationship)
Base R Alternative (No Packages Needed)
If you don't want to use the tidyverse, you can achieve the same result with base R using match() to map IDs to roles, then paste() to concatenate:
# Base R approach kinship_with_roles_base <- kinshipEstimates # Map ID1 and ID2 to their respective roles kinship_with_roles_base$Role1 <- sampleManifest$Role[match(kinship_with_roles_base$ID1, sampleManifest$SampleName)] kinship_with_roles_base$Role2 <- sampleManifest$Role[match(kinship_with_roles_base$ID2, sampleManifest$SampleName)] # Create the Roles column kinship_with_roles_base$Roles <- paste(kinship_with_roles_base$Role1, kinship_with_roles_base$Role2, sep = "-") # Reorder columns to match your desired output kinship_with_roles_base <- kinship_with_roles_base[, c("FID", "ID1", "ID2", "Roles", "Kinship", "Relationship")]
Verify the Result
Either approach will give you exactly the output you're looking for: a Roles column with the concatenated roles of each sample pair, formatted like sibling-proband, father-mother, etc.
内容的提问来源于stack exchange,提问作者Carmen Sandoval

