如何在R中基于联系人属性计算自我子网络的平均亲密度?
问题:计算自我中心网络中C类朋友间的平均亲密度
我有一份自我中心网络(ego-network)数据,参与者把自己的3位朋友分为A、B、C三类,同时给朋友间的亲密度打了分。现在想新增一列,计算每位参与者的C类朋友之间的平均亲密度。
比如第一行里,只有朋友1和朋友2都是C类,所以只用friend1_closeness_friend2的值计算,另外两个亲密度评分忽略。
数据示例
library(dplyr) data <- data.frame( ID = c("001", "002", "003"), friendshipType_1 = c("C", "C", "A"), friendshipType_2 = c("C", "C", "B"), friendshipType_3 = c("A", "C" , "A"), friend1_closeness_friend2 = c(1, 2, 3), friend1_closeness_friend3 = c(4, 3, 2), friend2_closeness_friend3 = c(1, 1, 2) )
期望输出
data <- data.frame( ID = c("001", "002", "003"), friendshipType_1 = c("C", "C", "A"), friendshipType_2 = c("C", "C", "B"), friendshipType_3 = c("A", "C" , "A"), friend1_closeness_friend2 = c(1, 2, 3), friend1_closeness_friend3 = c(4, 3, 2), friend2_closeness_friend3 = c(1, 1, 2), mean_c = c(1, 2, NA) )
解决方案
可以利用dplyr的行处理功能,逐行判断每对朋友是否都属于C类,筛选出符合条件的亲密度值后计算平均值:
方法一:简洁版
data <- data %>% rowwise() %>% mutate( mean_c = mean( c( # 判断朋友1和朋友2是否都是C类,是则取对应亲密度,否则设为NA ifelse(friendshipType_1 == "C" & friendshipType_2 == "C", friend1_closeness_friend2, NA), # 判断朋友1和朋友3是否都是C类 ifelse(friendshipType_1 == "C" & friendshipType_3 == "C", friend1_closeness_friend3, NA), # 判断朋友2和朋友3是否都是C类 ifelse(friendshipType_2 == "C" & friendshipType_3 == "C", friend2_closeness_friend3, NA) ), na.rm = TRUE ) ) %>% ungroup() %>% # 把全NA情况产生的NaN转为NA,匹配期望输出 mutate(mean_c = ifelse(is.nan(mean_c), NA, mean_c))
方法二:分步清晰版
如果希望逻辑更直观,可以先收集符合条件的亲密度值,再计算均值:
data <- data %>% rowwise() %>% mutate( # 收集所有C类朋友间的亲密度值,不符合条件的不加入列表 c_closeness = list( c( if (friendshipType_1 == "C" & friendshipType_2 == "C") friend1_closeness_friend2, if (friendshipType_1 == "C" & friendshipType_3 == "C") friend1_closeness_friend3, if (friendshipType_2 == "C" & friendshipType_3 == "C") friend2_closeness_friend3 ) ), # 计算列表中数值的均值,空列表时返回NA mean_c = if (length(c_closeness) == 0) NA else mean(unlist(c_closeness)) ) %>% # 移除临时中间列 select(-c_closeness) %>% ungroup()
两种方法都能得到与期望输出完全一致的结果。
内容的提问来源于stack exchange,提问作者Marie
相关产品推荐
相关产品推荐

