You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:高效统计分组元素在用户编码列表中的匹配人数(替代嵌套循环)

高效统计分组对应的有效用户数(R语言)

数据定义

分组数据框 df

df <- data.frame(
           groups=I(list(c("a"), c("b","c", "d", "e","f"), c("g","h"), c("i")))
)
df$group_count <- 0

用户编码数据框 people_codes

people_codes<-data.frame(
    eid=c(1,2,3,4,5,6, 7, 8, 9),
    present_code=I(list(c("g", "h"), c("a"), c("i"), c("g", "h"), c("a"), c("i", "a"), c("h"), c("f"), c("e"))))

需求说明

统计df中每个groups分组对应的有效用户数:只要某用户的present_code包含该分组中的至少一个元素,该用户即计入该分组的统计。

预期输出:

groups                 group_count
    "a"                     3
    "b","c", "d", "e","f"   2
    "g","h"                 3
    "i"                     2

问题:嵌套循环效率低下

原实现通过嵌套循环遍历两个数据框,当数据量较大时运行效率极低:

for (row1 in 1:nrow(df)) {
  for (row2 in 1:nrow(people_codes)){
    if (any(people_codes[row2, "present_code"][[1]] %in% df[row1, "groups"][[1]])){
      df[row1, "group_count"] <- df[row1, "group_count"]+1
    }
  }
}

高效解决方案

方法1:tidyverse工具链实现(推荐)

通过展开列表、关联匹配、分组统计的方式完成向量化操作,彻底规避显式循环:

library(tidyverse)

# 展开用户数据,生成"用户-编码"的唯一映射
people_expanded <- people_codes %>%
  unnest_longer(present_code) %>%
  distinct(eid, present_code)

# 展开分组数据,保留原分组的标识与标签
df_expanded <- df %>%
  select(groups) %>%
  unnest_longer(groups) %>%
  group_by(group_id = row_number()) %>%
  mutate(group_label = paste(groups, collapse = ", ")) %>%
  ungroup()

# 匹配用户与分组,统计每个分组的有效唯一用户数
result <- df_expanded %>%
  inner_join(people_expanded, by = c("groups" = "present_code")) %>%
  distinct(group_id, eid) %>%
  count(group_id, name = "group_count") %>%
  left_join(df_expanded %>% distinct(group_id, group_label), by = "group_id") %>%
  select(groups = group_label, group_count) %>%
  arrange(match(groups, sapply(df$groups, paste, collapse = ", ")))

print(result)

方法2:Base R向量化实现

利用矩阵运算与any的向量化特性,避免嵌套循环:

# 获取所有唯一编码
all_codes <- unique(unlist(c(df$groups, people_codes$present_code)))

# 生成用户-编码的匹配逻辑矩阵
user_match_mat <- sapply(all_codes, function(code) {
  sapply(people_codes$present_code, function(x) code %in% x)
})

# 生成分组-编码的包含逻辑矩阵
group_match_mat <- sapply(all_codes, function(code) {
  sapply(df$groups, function(x) code %in% x)
})

# 计算每个分组的有效用户数:用户与分组存在任意编码匹配则计数
df$group_count <- colSums(group_match_mat %*% t(user_match_mat) > 0)

print(df)

内容的提问来源于stack exchange,提问作者Caterina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 08:50:32