在R中基于现有数据框观测关联创建新数据框的实现需求
Solution to Generate Linked Individuals Data Frame
Here's a straightforward approach using dplyr and tidyr to create your desired Df.Result data frame. This solution groups individuals by their institution, generates all possible cross-pairs within each group, and formats the output to match your requested structure.
Step 1: Reproduce the Sample Input
First, let's create your input data frame for testing:
# Sample input data frame Df.1 <- data.frame( "Ind. Name" = c("J. Smith", "K. Kapplan", "A. Lindt", "B. Johnson", "E. Pitt", "S. Mathews", "P. Rawles", "P. Right", "O. Stray"), "Ind. ID" = c(12345, 12346, 12347, 12348, 12349, 12351, 12351, 12352, 12353), "Inst" = c("A", "A", "A", "B", "B", "C", "C", "C", "C"), "Inst. ID" = c(532, 532, 532, 761, 761, 890, 890, 890, 890), stringsAsFactors = FALSE )
Step 2: Define the Transformation Function
This function will handle grouping, pair generation, and formatting:
library(dplyr) library(tidyr) generate_linked_individuals <- function(input_df) { input_df %>% # Add a unique row ID per group to handle duplicate Ind.IDs group_by(Inst, `Inst. ID`) %>% mutate(row_id = row_number()) %>% ungroup() %>% # Generate all possible pairs of row IDs within each institution group group_by(Inst, `Inst. ID`) %>% expand(row_id, linked_row_id = row_id) %>% # Exclude self-pairs (no one linked to themselves) filter(row_id != linked_row_id) %>% # Join back to get individual details for both members of the pair left_join(input_df %>% mutate(row_id = row_number()), by = c("Inst", "Inst. ID", "row_id")) %>% left_join(input_df %>% mutate(row_id = row_number()), by = c("Inst", "Inst. ID", "linked_row_id" = "row_id"), suffix = c("", ".linked")) %>% # Rename and select columns to match your desired output select(`Ind. Name`, `Ind. ID`, `Linked Ind.` = `Ind. Name.linked`, `Linked Ind.ID` = `Ind. ID.linked`, Inst, `Inst. ID`) %>% ungroup() }
Step 3: Generate the Result
Call the function on your input data frame:
Df.Result <- generate_linked_individuals(Df.1)
Key Notes:
- Handling Duplicate IDs: The function uses a
row_idto distinguish individuals even if they share the sameInd. ID(like S. Mathews and P. Rawles in your sample). - Including Reverse Pairs: By default, this generates both directions of each pair (e.g., J. Smith → K. Kapplan and K. Kapplan → J. Smith). If you only want unique, non-reverse pairs, modify the filter to:
filter(row_id < linked_row_id)
Example Output Snippet
The resulting Df.Result will look like this (showing the first few rows):
Ind. Name Ind. ID Linked Ind. Linked Ind.ID Inst Inst. ID 1 J. Smith 12345 K. Kapplan 12346 A 532 2 J. Smith 12345 A. Lindt 12347 A 532 3 K. Kapplan 12346 J. Smith 12345 A 532 4 K. Kapplan 12346 A. Lindt 12347 A 532 5 A. Lindt 12347 J. Smith 12345 A 532 6 A. Lindt 12347 K. Kapplan 12346 A 532
内容的提问来源于stack exchange,提问作者Elli
相关产品推荐
相关产品推荐

