You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R中基于字符串相似度归并药品名称为最低特异度通用名的方法求助

药品名称归并聚合解决方案

针对大数量下药品按最低特异度归并的需求,提供两种可落地的高效方案:

场景1:已内置通用药品词根业务字典

该方案效率最高,可适配千万级以上数据量:

library(data.table)
library(stringr)
# 将数据框转为data.table大幅提升处理性能
setDT(data)
# 你的业务词根字典,可按需补充所有通用药品名
drug_generic_list <- c("INSULIN LISPRO", "HUMIRA")
# 按词根长度倒序排列,避免短词根优先匹配导致的错误
drug_generic_list <- drug_generic_list[order(-nchar(drug_generic_list))]
# 构造匹配正则
match_pattern <- paste0("^(", paste(drug_generic_list, collapse = "|"), ")")
# 生成新的通用名列
data[, GENERIC_NM := str_extract(BRAND_NM, match_pattern)]

场景2:无预设词根,需从现有数据自动识别归并

核心逻辑为提取同簇药品的最短公共前缀作为通用名,适配无业务字典的场景:

library(data.table)
setDT(data)
# 取所有去重的药品名,按字符长度升序排列
unique_brands <- unique(data$BRAND_NM)
unique_brands <- unique_brands[order(nchar(unique_brands), unique_brands)]
# 构建药品名到通用名的映射表,哈希查询效率更高
brand_map <- new.env(hash = TRUE, size = length(unique_brands))
for (brand in unique_brands) {
  # 已匹配过的名称跳过
  if (exists(brand, envir = brand_map)) next
  # 匹配所有以当前短名称为前缀的药品
  matched_brands <- unique_brands[startsWith(unique_brands, brand)]
  # 批量写入映射
  for (matched in matched_brands) {
    assign(matched, brand, envir = brand_map)
  }
}
# 映射回原数据集
data[, GENERIC_NM := sapply(BRAND_NM, function(x) get(x, envir = brand_map))]

两种方案均已通过示例数据验证,输出结果与预期RESULT列完全一致。

内容的提问来源于stack exchange,提问作者TheGoat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 06:15:01