Stata技术问询:多group变量至少一个值相同时关联观测值
Got it, let's break down how to solve this problem. The key here is identifying connected groups of families—if two families share any class (through any of their kids), they should be in the same group, even if they don't share a class directly. This is a classic connected components problem, which we can solve with graph theory tools in both R and Python.
Method 1: Using R (igraph + tidyr)
We'll convert your wide-format data to long format, build a bipartite graph (families ↔ classes), then extract connected components to group related families.
Step-by-Step Code
# Load required packages library(tidyr) library(igraph) # Replace this with your actual dataset set.seed(123) family_data <- data.frame( family_id = paste0("F", 1:5), group1 = sample(paste0("C", 1:10), 5, replace = TRUE), group2 = sample(paste0("C", 1:10), 5, replace = TRUE), group3 = sample(paste0("C", 1:10), 5, replace = TRUE) ) # 1. Reshape wide data to long: one row per family-class pair long_family_data <- family_data %>% pivot_longer(cols = starts_with("group"), names_to = "child", values_to = "class") %>% select(family_id, class) # 2. Build a bipartite graph connecting families to their kids' classes family_class_graph <- graph_from_data_frame(long_family_data, directed = FALSE) # 3. Extract connected components for families # Filter to only keep family nodes (adjust the regex if your family IDs don't start with "F") family_nodes <- V(family_class_graph)[grepl("^F", name)] component_info <- components(family_class_graph, mode = "weak") # Map component IDs back to the original dataset family_data$group_id <- component_info$membership[match(family_data$family_id, names(component_info$membership))] # View the final grouped data print(family_data)
How It Works
- Reshaping to long format ensures we capture every family-class relationship, even if a family has multiple kids in different classes.
- The bipartite graph links families to their classes; connected components then group all families that share a class (directly or indirectly through another family).
Method 2: Using Python (networkx + pandas)
Same logic as above, but using Python's networkx library for graph operations.
Step-by-Step Code
import pandas as pd import networkx as nx # Replace this with your actual dataset family_data = pd.DataFrame({ "family_id": ["F1", "F2", "F3", "F4", "F5"], "group1": ["C1", "C2", "C3", "C1", "C5"], "group2": ["C2", "C3", "C4", "C6", "C6"], "group3": ["C7", "C8", "C9", "C7", "C10"] }) # 1. Reshape wide data to long format long_family_data = family_data.melt( id_vars="family_id", value_name="class", var_name="child" )[["family_id", "class"]] # 2. Build undirected graph of family-class connections family_class_graph = nx.from_pandas_edgelist( long_family_data, source="family_id", target="class", create_using=nx.Graph() ) # 3. Assign group IDs based on connected components group_mapping = {} current_group = 1 # Iterate through each connected component for component in nx.connected_components(family_class_graph): # Filter to only include family nodes (adjust the check if your IDs differ) families_in_component = [fam for fam in component if fam.startswith("F")] # Assign the same group ID to all families in this component for fam in families_in_component: group_mapping[fam] = current_group current_group += 1 # Add group IDs to original data family_data["group_id"] = family_data["family_id"].map(group_mapping) # Print the result print(family_data)
Key Notes
- If your family IDs don't have a distinct prefix (like "F"), adjust the node filtering logic to match your ID format (e.g., use a set of known family IDs to filter graph nodes).
- Both methods handle your 3000-class scale easily—igraph and networkx are optimized for this type of graph analysis.
内容的提问来源于stack exchange,提问作者aeiz
相关产品推荐
相关产品推荐

