网络分析新手求助:从事件级数据构建完整组织关联网络
Hey there! I see exactly where you're stuck—your current setup only maps connections to Jobbik, but we need to capture every pair of organizations that co-hosted an event together, regardless of whether Jobbik is involved. Let's walk through how to adjust your workflow to make that happen.
The Core Issue
Your original code counts how many times each organization appeared with Jobbik, but we need to count how many times any two organizations showed up in the same event. This requires generating all unique pairs of organizations per event, then tallying those pairs across all events.
Step-by-Step Solution
Let's build on the cleaned long-format data you already have, with a few key tweaks:
1. First, Clean Up the Long-Format Data (Optional but Recommended)
First, let's make sure we have unique organizations per event (no duplicates) and remove any empty entries:
# Start with your cleaned long data jobbik_clean <- jobbik %>% # Remove any empty organization names filter(org_names != "") %>% # For each event, keep only unique organizations (avoid duplicate pairs later) group_by(id) %>% distinct(org_names, .keep_all = TRUE) %>% ungroup()
2. Generate All Organizational Pairs Per Event
This is the critical step we were missing. For each event, we'll create every possible unique pair of organizations that participated:
library(purrr) # We'll use this for mapping combn() over events # Create all unique pairs of organizations per event org_pairs <- jobbik_clean %>% group_by(id) %>% # For each event, generate all 2-organization combinations mutate(pairs = map(org_names, ~ combn(org_names, 2, simplify = FALSE))) %>% # Unnest the list of pairs into individual rows unnest(pairs) %>% # Split each pair into "from" and "to", and sort them so A-B = B-A mutate( from = map_chr(pairs, ~ sort(.x)[1]), to = map_chr(pairs, ~ sort(.x)[2]) ) %>% # Keep only the columns we need select(id, from, to) %>% ungroup()
Sorting from and to ensures we don't count the same pair twice (e.g., "Org A-Jobbik" and "Jobbik-Org A" are treated as one pair).
3. Calculate Edge Weights (Co-Occurrence Count)
Now we'll count how many times each pair of organizations appeared together across all events—this becomes our edge weight:
edges <- org_pairs %>% # Count occurrences of each from-to pair, name the count "weight" count(from, to, name = "weight")
4. Build the Network & Plot
We can use the edges data frame directly to build our graph (no need for separate node IDs unless you prefer them—igraph works fine with character labels):
# Create the undirected graph routes_igraph <- graph_from_data_frame(d = edges, directed = FALSE) # Plot with helpful visual cues plot(routes_igraph, layout = layout_with_graphopt, edge.width = E(routes_igraph)$weight / 2, # Thicker edges = more co-occurrences vertex.size = degree(routes_igraph) * 2, # Larger nodes = more connections vertex.label.cex = 0.7, # Smaller labels for readability vertex.color = ifelse(V(routes_igraph)$name == "Jobbik Magyarországért Mozgalom", "red", "lightblue") # Highlight Jobbik )
What This Does Differently
- Instead of only linking organizations to Jobbik, this network shows all connections between organizations that co-hosted events.
- Edge weights represent how often two organizations worked together, so stronger partnerships are visually clearer.
- The plot highlights Jobbik (in red) so you can still easily track its connections alongside other organizational links.
内容的提问来源于stack exchange,提问作者Erdne Htábrob

