R语言中子串位置定位与出现顺序排名及数据框存储实现
R语言中子串出现顺序排名与数据框生成方案
需求描述
在R语言中查找指定子串在长字符串中的出现位置,按出现先后顺序生成排名,最终将结果存入数据框用于后续重排操作。
示例场景
现有3个长字符串,需确定子串PR、FR、VG的出现顺序,将各子串的出现顺序排名存入数据框。
输入代码
fullStrings <- c(' FR Banana VG Carrot VG Celery PR Chicken ', ' VG Broccoli PR Tofu VG Celery FR Apple ', ' PR Pork VG Brussels Sprouts FR Orange VG Carrot ') findStrings <- c(' PR ', ' FR ', ' VG ')
用户问题
尝试使用str_locate_all函数解析输出,但无法得到期望格式的结果,期望生成如下结构的数据框:
# 期望结果示例构造 string <- c(fullStrings[1], fullStrings[2], fullStrings[3]) PRpos <- c(4, 2, 1) FRpos <- c(1, 4, 3) VG1pos <- c(2, 1, 2) VG2pos <- c(3, 3, 4) foodLocations <- data.frame(string, PRpos, FRpos, VG1pos, VG2pos) foodLocations # 输出: # string PRpos FRpos VG1pos VG2pos # 1 FR Banana VG Carrot VG Celery PR Chicken 4 1 2 3 # 2 VG Broccoli PR Tofu VG Celery FR Apple 2 4 1 3 # 3 PR Pork VG Brussels Sprouts FR Orange VG Carrot 1 3 2 4
解决方案
使用stringr、purrr和dplyr包组合处理,以下是完整代码:
library(stringr) library(purrr) library(dplyr) # 输入数据 fullStrings <- c(' FR Banana VG Carrot VG Celery PR Chicken ', ' VG Broccoli PR Tofu VG Celery FR Apple ', ' PR Pork VG Brussels Sprouts FR Orange VG Carrot ') findStrings <- c(' PR ', ' FR ', ' VG ') # 定义处理单个字符串的函数,返回子串排名列表 get_substring_ranks <- function(str, patterns) { # 收集所有子串的出现位置 all_matches <- map_dfr(patterns, function(p) { locs <- str_locate_all(str, p)[[1]][, 1] if (length(locs) == 0) { return(tibble(pos = numeric(), pattern = p)) } tibble(pos = locs, pattern = p) }) # 按出现位置排序,生成排名 all_matches <- all_matches %>% arrange(pos) %>% mutate(rank = row_number()) # 整理为目标格式:多出现的子串添加编号后缀 result_list <- list() for (p in patterns) { pattern_matches <- all_matches %>% filter(pattern == p) pattern_name <- str_remove_all(p, " ") if (nrow(pattern_matches) == 0) { result_list[[pattern_name]] <- NA } else { for (i in seq_len(nrow(pattern_matches))) { col_name <- paste0(pattern_name, i, "pos") result_list[[col_name]] <- pattern_matches$rank[i] } } } return(result_list) } # 批量处理所有字符串,生成最终数据框 foodLocations <- map_dfr(fullStrings, get_substring_ranks, patterns = findStrings, .id = "string_idx") %>% mutate(string = fullStrings[as.integer(string_idx)]) %>% select(string, everything(), -string_idx) # 查看结果 print(foodLocations)
代码说明
- 收集子串位置:遍历所有目标子串,用
str_locate_all获取每个子串在长字符串中的起始位置,存入临时数据框。 - 生成排名:按子串出现的位置排序,为每个子串分配先后排名。
- 整理格式:对每个子串,根据出现次数生成带编号的列名(如
VG出现两次则生成VG1pos、VG2pos),并填充对应排名。 - 合并结果:批量处理所有长字符串,合并为最终的数据框,保留原字符串列。
运行代码后即可得到与期望一致的结果。
内容的提问来源于stack exchange,提问作者CoolGuyHasChillDay
相关产品推荐
相关产品推荐

