You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Stata技术问询:多group变量至少一个值相同时关联观测值

Solution: Group Families by Shared Class Connections

Got it, let's break down how to solve this problem. The key here is identifying connected groups of families—if two families share any class (through any of their kids), they should be in the same group, even if they don't share a class directly. This is a classic connected components problem, which we can solve with graph theory tools in both R and Python.


Method 1: Using R (igraph + tidyr)

We'll convert your wide-format data to long format, build a bipartite graph (families ↔ classes), then extract connected components to group related families.

Step-by-Step Code

# Load required packages
library(tidyr)
library(igraph)

# Replace this with your actual dataset
set.seed(123)
family_data <- data.frame(
  family_id = paste0("F", 1:5),
  group1 = sample(paste0("C", 1:10), 5, replace = TRUE),
  group2 = sample(paste0("C", 1:10), 5, replace = TRUE),
  group3 = sample(paste0("C", 1:10), 5, replace = TRUE)
)

# 1. Reshape wide data to long: one row per family-class pair
long_family_data <- family_data %>%
  pivot_longer(cols = starts_with("group"), names_to = "child", values_to = "class") %>%
  select(family_id, class)

# 2. Build a bipartite graph connecting families to their kids' classes
family_class_graph <- graph_from_data_frame(long_family_data, directed = FALSE)

# 3. Extract connected components for families
# Filter to only keep family nodes (adjust the regex if your family IDs don't start with "F")
family_nodes <- V(family_class_graph)[grepl("^F", name)]
component_info <- components(family_class_graph, mode = "weak")

# Map component IDs back to the original dataset
family_data$group_id <- component_info$membership[match(family_data$family_id, names(component_info$membership))]

# View the final grouped data
print(family_data)

How It Works

  • Reshaping to long format ensures we capture every family-class relationship, even if a family has multiple kids in different classes.
  • The bipartite graph links families to their classes; connected components then group all families that share a class (directly or indirectly through another family).

Method 2: Using Python (networkx + pandas)

Same logic as above, but using Python's networkx library for graph operations.

Step-by-Step Code

import pandas as pd
import networkx as nx

# Replace this with your actual dataset
family_data = pd.DataFrame({
    "family_id": ["F1", "F2", "F3", "F4", "F5"],
    "group1": ["C1", "C2", "C3", "C1", "C5"],
    "group2": ["C2", "C3", "C4", "C6", "C6"],
    "group3": ["C7", "C8", "C9", "C7", "C10"]
})

# 1. Reshape wide data to long format
long_family_data = family_data.melt(
    id_vars="family_id", 
    value_name="class", 
    var_name="child"
)[["family_id", "class"]]

# 2. Build undirected graph of family-class connections
family_class_graph = nx.from_pandas_edgelist(
    long_family_data, 
    source="family_id", 
    target="class", 
    create_using=nx.Graph()
)

# 3. Assign group IDs based on connected components
group_mapping = {}
current_group = 1

# Iterate through each connected component
for component in nx.connected_components(family_class_graph):
    # Filter to only include family nodes (adjust the check if your IDs differ)
    families_in_component = [fam for fam in component if fam.startswith("F")]
    # Assign the same group ID to all families in this component
    for fam in families_in_component:
        group_mapping[fam] = current_group
    current_group += 1

# Add group IDs to original data
family_data["group_id"] = family_data["family_id"].map(group_mapping)

# Print the result
print(family_data)

Key Notes

  • If your family IDs don't have a distinct prefix (like "F"), adjust the node filtering logic to match your ID format (e.g., use a set of known family IDs to filter graph nodes).
  • Both methods handle your 3000-class scale easily—igraph and networkx are optimized for this type of graph analysis.

内容的提问来源于stack exchange,提问作者aeiz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:06:06