基于R语言DataFrame的图节点wtarget关联次数统计需求
图节点标签关联次数统计解决方案
嘿,我来帮你搞定这个图节点的标签关联统计问题!先看你的原始数据:
df <- data.frame( source = c("a","a","a",'z1','b'), target = c("b","c","d",'a','e'), wsource = c('w1','w2','w1','w2','w1'), wtarget = c('w1','w1','w1','w1','w2') )
对应的表格是:
| source | target | wsource | wtarget |
|---|---|---|---|
| a | b | w1 | w1 |
| a | c | w2 | w1 |
| a | d | w1 | w1 |
| z1 | a | w2 | w1 |
| b | e | w1 | w2 |
我猜你是要统计每个节点的两类次数之和:
- 当节点作为
source时,所有关联边中wtarget的出现总次数(或按标签分组的次数) - 当节点作为
target时,所有关联边中wsource的出现总次数
下面给你几种实用的解决方案:
方案1:统计每个节点的总关联次数(不区分标签)
如果你只需要每个节点的总次数(不管标签类型,只算关联的边数对应的标签次数),用dplyr处理最顺手:
首先加载包(没装的话先跑install.packages("dplyr")):
library(dplyr)
然后分三步处理:
# 1. 统计每个source节点对应的wtarget总次数(每条边算1次) source_counts <- df %>% group_by(node = source) %>% summarise(source_wtarget_count = n(), .groups = "drop") # 2. 统计每个target节点对应的wsource总次数(每条边算1次) target_counts <- df %>% group_by(node = target) %>% summarise(target_wsource_count = n(), .groups = "drop") # 3. 合并结果,填充缺失值为0,计算总和 final_counts <- full_join(source_counts, target_counts, by = "node") %>% mutate( source_wtarget_count = replace_na(source_wtarget_count, 0), target_wsource_count = replace_na(target_wsource_count, 0), total_count = source_wtarget_count + target_wsource_count ) # 查看结果 print(final_counts)
运行后会得到:
# A tibble: 6 × 4 node source_wtarget_count target_wsource_count total_count <chr> <int> <int> <int> 1 a 3 1 4 2 b 1 1 2 3 z1 1 0 1 4 c 0 1 1 5 d 0 1 1 6 e 0 1 1
举个例子:节点a作为source出现3次(对应3次wtarget),作为target出现1次(对应1次wsource),总和就是4,完美符合预期。
方案2:按标签分组统计每个节点的关联次数
如果需要按标签细分(比如知道节点a关联的wtarget是w1的次数有多少,关联的wsource是w2的次数有多少),可以这样写:
# 1. 统计source节点的wtarget标签分布 source_tag_counts <- df %>% group_by(node = source, tag = wtarget) %>% summarise(count = n(), .groups = "drop") %>% mutate(type = "source_wtarget") # 2. 统计target节点的wsource标签分布 target_tag_counts <- df %>% group_by(node = target, tag = wsource) %>% summarise(count = n(), .groups = "drop") %>% mutate(type = "target_wsource") # 3. 合并所有结果并排序 final_tag_counts <- bind_rows(source_tag_counts, target_tag_counts) %>% arrange(node, type, tag) # 查看结果 print(final_tag_counts)
运行结果会清晰展示每个节点对应不同标签的次数:
# A tibble: 8 × 4 node tag type count <chr> <chr> <chr> <int> 1 a w1 source_wtarget 3 2 a w2 target_wsource 1 3 b w1 target_wsource 1 4 b w2 source_wtarget 1 5 c w1 target_wsource 1 6 d w1 target_wsource 1 7 e w1 target_wsource 1 8 z1 w1 source_wtarget 1
方案3:用基础R实现(不依赖第三方包)
要是你不想用tidyverse系列的包,纯基础R也能搞定总次数统计:
# 获取所有唯一节点 all_nodes <- unique(c(df$source, df$target)) # 初始化结果数据框 result <- data.frame( node = all_nodes, source_wtarget_count = 0, target_wsource_count = 0, total_count = 0, stringsAsFactors = FALSE ) # 统计source节点的wtarget总次数 source_tab <- table(df$source) result$source_wtarget_count[match(names(source_tab), result$node)] <- as.numeric(source_tab) # 统计target节点的wsource总次数 target_tab <- table(df$target) result$target_wsource_count[match(names(target_tab), result$node)] <- as.numeric(target_tab) # 计算总和 result$total_count <- result$source_wtarget_count + result$target_wsource_count # 查看结果 print(result)
这个结果和方案1完全一致,适合不想额外装包的场景。
内容的提问来源于stack exchange,提问作者nhern121
相关产品推荐
相关产品推荐

