You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中向量化嵌套循环?ICD-10代码匹配计数优化

向量化嵌套循环统计ICD-10代码匹配次数

问题描述

需要统计full_df每行的coding_19列表中,至少有一个代码出现在icd10_codes每个人员的present_icd10列表中的次数,替代原本的嵌套循环实现。

原始嵌套循环代码

for (row1 in 1:nrow(full_df)) {
  for (row2 in 1:nrow(icd10_codes)){
    if (any(full_df[row1, "coding_19"] %in% icd10_codes[row2, "present_icd10"])){
      full_df[row1, "code_count"] <- full_df[row1, "code_count"]+1
    }
  }
}

数据定义

full_df

full_df <- data.frame(
           coding_19=I(list(c("H353"), c("B20","B21", "B22", "B23","B24","Z21", "F024","O987"), c("G30","F00"), c("E780")))
)
full_df$code_count <- 0

icd10_codes

icd10_codes <- data.frame(
    eid=c(1,2,3,4,5,6),
    present_icd10=I(list(c("G30", "F00"), c("B20"), c("E780"), c("H401", "H409"), c("H353"), c("E780")))
)

预期输出

coding_19                                              code_count
"H353"                                                     1
"B20", "B21", "B22", "B23", "B24", "Z21", "F024", "O987"   1
"G30", "F00"                                               1
"E780"                                                     2

尝试的无效代码

full_df <- full_df %>%
  rowwise() %>%
  mutate(code_count = code_count + as.integer(any(coding_19 %in% icd10_codes$present_icd10)))

解决方案

方法1:使用purrr包的向量化映射

利用map_int遍历full_df的每行coding_19,再对每个人员的present_icd10检查交集,最后求和:

library(purrr)

full_df$code_count <- map_int(full_df$coding_19, function(codes) {
  sum(map_lgl(icd10_codes$present_icd10, ~ any(.x %in% codes)))
})

方法2:结合tidyverse的行处理

通过rowwise配合map_lgl实现逐行统计:

library(dplyr)
library(purrr)

full_df <- full_df %>%
  rowwise() %>%
  mutate(code_count = sum(map_lgl(icd10_codes$present_icd10, ~ any(coding_19 %in% .x)))) %>%
  ungroup()

方法3:基础R的向量化实现

不依赖第三方包,用sapply嵌套完成统计:

full_df$code_count <- sapply(full_df$coding_19, function(codes) {
  sum(sapply(icd10_codes$present_icd10, function(present) any(present %in% codes)))
})

原理说明

以上方法均将嵌套循环转化为向量化遍历操作:

  1. 外层遍历full_df的每个coding_19列表
  2. 内层遍历icd10_codes的每个present_icd10列表,通过any(x %in% y)判断两个列表是否存在交集
  3. 将内层的逻辑结果(TRUE/FALSE)自动转为数值(1/0)后求和,得到每个coding_19对应的匹配人员数

执行后输出结果与预期完全一致:

print(full_df)
#                                            coding_19 code_count
# 1                                              H353          1
# 2 B20, B21, B22, B23, B24, Z21, F024, O987          1
# 3                                          G30, F00          1
# 4                                              E780          2

内容的提问来源于stack exchange,提问作者Caterina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 16:34:52