You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言的iGraph中识别有向网络中的双向关注匹配

Hey Kimberly, let's break this down step by step for your Twitter directed graph analysis using iGraph—this is a super common use case, so we've got solid, actionable patterns to follow.

核心分析步骤与代码实现

1. 识别双向关注(互粉)关系

In a directed graph, mutual follows mean there's both an edge from User A to User B and from User B to User A. iGraph has built-in tools to streamline this, but we can also make explicit checks for clarity:

import igraph as ig

# 假设你已经构建好有向图,命名为g
# 方法1:用iGraph内置的互边标记
mutual_edges = g.es.select(_is_mutual=True)
# 提取去重的互粉节点对(避免重复记录(A,B)和(B,A))
mutual_node_pairs = [(e.source, e.target) for e in mutual_edges if e.source < e.target]

# 方法2:手动验证反向边(适合理解逻辑)
mutual_pairs = set()
for edge in g.es:
    source, target = edge.source, edge.target
    if g.are_connected(target, source):
        # 用sorted确保每对只存一次
        pair = tuple(sorted((source, target)))
        mutual_pairs.add(pair)

2. 筛选未获得回关的用户

We can split this into two groups: users who follow others but aren't followed back, and users who are followed but don't follow back. Here's how to extract both:

# 获取所有单向关注的边(A→B存在,但B→A不存在)
unidirectional_edges = g.es.select(_is_mutual=False)

# 统计两类用户
unfollowed_followers = set()  # 关注了别人却没被回关的用户
unfollowing_targets = set()   # 被别人关注却没回关的用户

for edge in unidirectional_edges:
    a, b = edge.source, edge.target
    unfollowed_followers.add(a)
    unfollowing_targets.add(b)

# 关联你的用户性别数据(假设user_data是{节点ID: {'gender': 'male/female/other'}})
unfollowed_male_count = len([uid for uid in unfollowed_followers if user_data[uid]['gender'] == 'male'])
unfollowed_female_count = len([uid for uid in unfollowed_followers if user_data[uid]['gender'] == 'female'])

3. 分析性别与回关率的关联

To test if male users get more follow-backs, we'll calculate follow-back rates (mutual follows ÷ total followers) for each gender, then run basic stats to check for significance:

import numpy as np
from scipy.stats import ttest_ind

# 计算每个用户的回关率
followback_rates = {}
for node in g.vs:
    node_id = node.index
    total_followers = node.indegree()
    if total_followers == 0:
        continue  # 跳过无粉丝的用户
    
    # 统计该用户回关了多少粉丝
    mutual_count = 0
    for follower in g.neighbors(node_id, mode='in'):
        if g.are_connected(node_id, follower):
            mutual_count += 1
    
    followback_rates[node_id] = mutual_count / total_followers

# 按性别分组计算平均回关率
male_rates = [rate for uid, rate in followback_rates.items() if user_data[uid]['gender'] == 'male']
female_rates = [rate for uid, rate in followback_rates.items() if user_data[uid]['gender'] == 'female']

male_avg = np.mean(male_rates) if male_rates else 0
female_avg = np.mean(female_rates) if female_rates else 0

print(f"男性用户平均回关率: {male_avg:.2f}")
print(f"女性用户平均回关率: {female_avg:.2f}")

# 做显著性检验(验证差异是否偶然)
if len(male_rates) > 1 and len(female_rates) > 1:
    stat, p_value = ttest_ind(male_rates, female_rates)
    print(f"T检验统计量: {stat:.2f}, P值: {p_value:.4f}")
    print("差异具有统计学显著性" if p_value < 0.05 else "差异无统计学显著性")

关键注意事项

  • 节点ID映射: 确保你的iGraph节点ID和用户数据表的ID完全对应,避免数据错位
  • 缺失数据: 单独处理性别未知的用户(可以排除或单独分组,避免干扰结果)
  • 性能优化: 对于超大型数据集,用iGraph的get_adjacency()矩阵代替循环遍历,会大幅提升速度

Let me know if you need help tweaking this to fit your specific dataset structure or if you hit performance snags with large graphs!

内容的提问来源于stack exchange,提问作者Kimberly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:11:09