在R语言的iGraph中识别有向网络中的双向关注匹配
Hey Kimberly, let's break this down step by step for your Twitter directed graph analysis using iGraph—this is a super common use case, so we've got solid, actionable patterns to follow.
1. 识别双向关注(互粉)关系
In a directed graph, mutual follows mean there's both an edge from User A to User B and from User B to User A. iGraph has built-in tools to streamline this, but we can also make explicit checks for clarity:
import igraph as ig # 假设你已经构建好有向图,命名为g # 方法1:用iGraph内置的互边标记 mutual_edges = g.es.select(_is_mutual=True) # 提取去重的互粉节点对(避免重复记录(A,B)和(B,A)) mutual_node_pairs = [(e.source, e.target) for e in mutual_edges if e.source < e.target] # 方法2:手动验证反向边(适合理解逻辑) mutual_pairs = set() for edge in g.es: source, target = edge.source, edge.target if g.are_connected(target, source): # 用sorted确保每对只存一次 pair = tuple(sorted((source, target))) mutual_pairs.add(pair)
2. 筛选未获得回关的用户
We can split this into two groups: users who follow others but aren't followed back, and users who are followed but don't follow back. Here's how to extract both:
# 获取所有单向关注的边(A→B存在,但B→A不存在) unidirectional_edges = g.es.select(_is_mutual=False) # 统计两类用户 unfollowed_followers = set() # 关注了别人却没被回关的用户 unfollowing_targets = set() # 被别人关注却没回关的用户 for edge in unidirectional_edges: a, b = edge.source, edge.target unfollowed_followers.add(a) unfollowing_targets.add(b) # 关联你的用户性别数据(假设user_data是{节点ID: {'gender': 'male/female/other'}}) unfollowed_male_count = len([uid for uid in unfollowed_followers if user_data[uid]['gender'] == 'male']) unfollowed_female_count = len([uid for uid in unfollowed_followers if user_data[uid]['gender'] == 'female'])
3. 分析性别与回关率的关联
To test if male users get more follow-backs, we'll calculate follow-back rates (mutual follows ÷ total followers) for each gender, then run basic stats to check for significance:
import numpy as np from scipy.stats import ttest_ind # 计算每个用户的回关率 followback_rates = {} for node in g.vs: node_id = node.index total_followers = node.indegree() if total_followers == 0: continue # 跳过无粉丝的用户 # 统计该用户回关了多少粉丝 mutual_count = 0 for follower in g.neighbors(node_id, mode='in'): if g.are_connected(node_id, follower): mutual_count += 1 followback_rates[node_id] = mutual_count / total_followers # 按性别分组计算平均回关率 male_rates = [rate for uid, rate in followback_rates.items() if user_data[uid]['gender'] == 'male'] female_rates = [rate for uid, rate in followback_rates.items() if user_data[uid]['gender'] == 'female'] male_avg = np.mean(male_rates) if male_rates else 0 female_avg = np.mean(female_rates) if female_rates else 0 print(f"男性用户平均回关率: {male_avg:.2f}") print(f"女性用户平均回关率: {female_avg:.2f}") # 做显著性检验(验证差异是否偶然) if len(male_rates) > 1 and len(female_rates) > 1: stat, p_value = ttest_ind(male_rates, female_rates) print(f"T检验统计量: {stat:.2f}, P值: {p_value:.4f}") print("差异具有统计学显著性" if p_value < 0.05 else "差异无统计学显著性")
关键注意事项
- 节点ID映射: 确保你的iGraph节点ID和用户数据表的ID完全对应,避免数据错位
- 缺失数据: 单独处理性别未知的用户(可以排除或单独分组,避免干扰结果)
- 性能优化: 对于超大型数据集,用iGraph的
get_adjacency()矩阵代替循环遍历,会大幅提升速度
Let me know if you need help tweaking this to fit your specific dataset structure or if you hit performance snags with large graphs!
内容的提问来源于stack exchange,提问作者Kimberly

