基于子串匹配实现Tibble中SectorId到ClusterId的映射
解决方案
首先加载所需的tidyverse包:
library(tidyverse)
步骤1:预处理映射表
给映射表添加行业代码长度列,并按长度降序排列,确保优先匹配更具体的层级(长前缀比短前缀优先级高):
map_tabl <- map_tabl %>% mutate(sector_len = nchar(SectorId)) %>% arrange(desc(sector_len))
步骤2:匹配ClusterId到股票表
逐行处理股票表,通过前缀匹配找到对应的ClusterId:
result_tabl <- tabl %>% rowwise() %>% mutate(ClusterId = map_tabl$ClusterId[str_detect(SectorId, paste0("^", map_tabl$SectorId))][1]) %>% ungroup()
验证结果
运行后输出的result_tabl即为目标表:
print(result_tabl) # # A tibble: 4 × 3 # Stock SectorId ClusterId # <chr> <chr> <chr> # 1 A 30101010 C1 # 2 B 30101010 C1 # 3 C 20103015 C2 # 4 D 55102010 C3
内容的提问来源于stack exchange,提问作者randomwalker
相关产品推荐
相关产品推荐

