R语言如何匹配两个数据框中斜杠分隔的标识列生成对应Id拼接列
问题核心
你之前的匹配失败原因是original的Identification列为多值拼接的字符串,无法直接和to_match的单值标识直接匹配,需要先拆分字段、逐值匹配后再拼接结果。
实现方法
方法1:基础R实现(无需安装额外包)
# 逐行拆分Identification、匹配、拼接 original$Id <- sapply(strsplit(original$Identification, "/"), function(x) { # 转成数值和to_match的Identification类型统一,避免匹配失败 match_ids <- as.numeric(x) paste(to_match$Id[match(match_ids, to_match$Identification)], collapse = "/") })
运行后original即可得到预期结果:
| Names | Identification | Id |
|---|---|---|
| Animals | 15/20/25/26 | Cat/Dog/Elephant/Mouse |
| Fruits | 1/2/3/4 | Banana/Melon/Mango/Apple |
方法2:tidyverse实现(语法更简洁)
library(dplyr) library(tidyr) original <- original %>% separate_rows(Identification, sep = "/", convert = TRUE) %>% # 拆分为长表,自动转换为数值类型 left_join(to_match, by = "Identification") %>% # 关联匹配Id group_by(Names) %>% summarise( Identification = paste(Identification, collapse = "/"), Id = paste(Id, collapse = "/") ) # 按分类分组拼接回原数据结构
内容的提问来源于stack exchange,提问作者CodingBiology
相关产品推荐
相关产品推荐

