You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用apply方法加速填充示例间Type交集矩阵?

问题描述

现有如下DataFrame:

Examples Type
1 example1    a
2 example1    b
3 example1    c
4 example1    c
5 example2    c

需求是构建一个行、列均对应唯一Examples值的矩阵,矩阵元素为对应两个Examples的Type集合的交集大小。

当前已初始化矩阵的代码:

my_mat <- matrix(0, nrow=length(unique(df$Examples)), ncol=length(unique(df$Examples))) 
rownames(my_mat) <- unique(df$Examples)
colnames(my_mat) <- unique(df$Examples)

但使用双重for循环填充矩阵时,数据量大时速度明显变慢:

get_intersection <- function(example1, example2) {
  return(length(dplyr::intersect(example1, example2)))
}

for (i in 1:nrow(my_mat)) {
  curr_row <- rownames(my_mat)[i]
  for (j in 1:ncol(my_mat)) {
    curr_col <- colnames(my_mat)[j]
    my_mat[i, j] <- get_intersection(df$Type[which(df$Examples %in% curr_row)], 
                                     df$Type[which(df$Examples %in% curr_col)])
  }
}

DataFrame的结构定义:

df <- structure(list(Examples = c("example1", "example1", "example1", 
"example1", "example2"), Type = c("a", "b", "c", "c", "c")), class = "data.frame", row.names = c(NA, 
-5L))

需要改用apply系列方法加速矩阵填充。

优化方案

核心思路

先预处理每个Examples对应的唯一Type集合,避免每次计算交集时重复从原DataFrame筛选数据;再利用apply系列函数的向量化特性替代双重for循环,提升运行效率。

具体实现

步骤1:预处理唯一Type集合

先提取每个Examples对应的去重Type列表:

# 获取所有唯一的Examples
unique_examples <- unique(df$Examples)
# 生成每个Examples对应的唯一Type集合列表
type_sets <- lapply(unique_examples, function(x) {
  unique(df$Type[df$Examples == x])
})
# 为列表命名,方便后续匹配
names(type_sets) <- unique_examples

步骤2:用outer+Vectorize生成矩阵

利用outer函数处理两两Examples的组合,结合Vectorize将交集计算函数向量化,直接生成目标矩阵:

# 定义交集大小计算函数
calc_intersect_size <- function(ex1, ex2) {
  length(intersect(type_sets[[ex1]], type_sets[[ex2]]))
}
# 向量化函数并生成矩阵
my_mat <- outer(unique_examples, unique_examples, Vectorize(calc_intersect_size))
# 设置行列名(可选,因为outer默认用unique_examples作为维度名)
rownames(my_mat) <- unique_examples
colnames(my_mat) <- unique_examples

步骤3:验证结果

运行后得到的矩阵如下:

example1 example2
example1        3        1
example2        1        1

符合预期:example1自身的Type交集是3种(a、b、c),example1与example2的交集是1种(c),example2自身交集是1种(c)。

替代方案:用apply遍历行

如果更习惯用apply,可以遍历矩阵的行,对每一行用sapply计算列对应的交集大小:

# 初始化矩阵(同原代码)
my_mat <- matrix(0, nrow=length(unique_examples), ncol=length(unique_examples))
rownames(my_mat) <- unique_examples
colnames(my_mat) <- unique_examples

# 用apply填充矩阵
my_mat <- t(apply(my_mat, 1, function(row_name) {
  sapply(colnames(my_mat), function(col_name) {
    length(intersect(type_sets[[row_name]], type_sets[[col_name]]))
  })
}))

效率提升原因

  1. 预处理去重:提前提取每个Examples的唯一Type集合,避免每次循环都从原DataFrame筛选、去重,减少重复计算。
  2. 向量化操作:apply系列函数(尤其是outer)底层用C语言实现循环,比R的显式双重for循环效率高得多,数据量越大,提升越明显。

内容的提问来源于stack exchange,提问作者user2657817

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 11:55:24