You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中子串位置定位与出现顺序排名及数据框存储实现

R语言中子串出现顺序排名与数据框生成方案

需求描述

在R语言中查找指定子串在长字符串中的出现位置,按出现先后顺序生成排名,最终将结果存入数据框用于后续重排操作。

示例场景

现有3个长字符串,需确定子串PR、FR、VG的出现顺序,将各子串的出现顺序排名存入数据框。

输入代码

fullStrings <- c(' FR Banana VG Carrot VG Celery PR Chicken ', ' VG Broccoli PR Tofu VG Celery FR Apple ', ' PR Pork VG Brussels Sprouts FR Orange VG Carrot ')
findStrings <- c(' PR ', ' FR ', ' VG ')

用户问题

尝试使用str_locate_all函数解析输出,但无法得到期望格式的结果,期望生成如下结构的数据框:

# 期望结果示例构造
string <- c(fullStrings[1], fullStrings[2], fullStrings[3])
PRpos <- c(4, 2, 1)
FRpos <- c(1, 4, 3)
VG1pos <- c(2, 1, 2)
VG2pos <- c(3, 3, 4)

foodLocations <- data.frame(string, PRpos, FRpos, VG1pos, VG2pos)
foodLocations
# 输出:
#                                              string PRpos FRpos VG1pos VG2pos
# 1         FR Banana VG Carrot VG Celery PR Chicken      4     1      2      3
# 2            VG Broccoli PR Tofu VG Celery FR Apple      2     4      1      3
# 3  PR Pork VG Brussels Sprouts FR Orange VG Carrot      1     3      2      4

解决方案

使用stringr、purrr和dplyr包组合处理,以下是完整代码:

library(stringr)
library(purrr)
library(dplyr)

# 输入数据
fullStrings <- c(' FR Banana VG Carrot VG Celery PR Chicken ', ' VG Broccoli PR Tofu VG Celery FR Apple ', ' PR Pork VG Brussels Sprouts FR Orange VG Carrot ')
findStrings <- c(' PR ', ' FR ', ' VG ')

# 定义处理单个字符串的函数,返回子串排名列表
get_substring_ranks <- function(str, patterns) {
  # 收集所有子串的出现位置
  all_matches <- map_dfr(patterns, function(p) {
    locs <- str_locate_all(str, p)[[1]][, 1]
    if (length(locs) == 0) {
      return(tibble(pos = numeric(), pattern = p))
    }
    tibble(pos = locs, pattern = p)
  })
  
  # 按出现位置排序,生成排名
  all_matches <- all_matches %>%
    arrange(pos) %>%
    mutate(rank = row_number())
  
  # 整理为目标格式:多出现的子串添加编号后缀
  result_list <- list()
  for (p in patterns) {
    pattern_matches <- all_matches %>% filter(pattern == p)
    pattern_name <- str_remove_all(p, " ")
    if (nrow(pattern_matches) == 0) {
      result_list[[pattern_name]] <- NA
    } else {
      for (i in seq_len(nrow(pattern_matches))) {
        col_name <- paste0(pattern_name, i, "pos")
        result_list[[col_name]] <- pattern_matches$rank[i]
      }
    }
  }
  return(result_list)
}

# 批量处理所有字符串,生成最终数据框
foodLocations <- map_dfr(fullStrings, get_substring_ranks, patterns = findStrings, .id = "string_idx") %>%
  mutate(string = fullStrings[as.integer(string_idx)]) %>%
  select(string, everything(), -string_idx)

# 查看结果
print(foodLocations)

代码说明

  1. 收集子串位置:遍历所有目标子串,用str_locate_all获取每个子串在长字符串中的起始位置,存入临时数据框。
  2. 生成排名:按子串出现的位置排序,为每个子串分配先后排名。
  3. 整理格式:对每个子串,根据出现次数生成带编号的列名(如VG出现两次则生成VG1pos、VG2pos),并填充对应排名。
  4. 合并结果:批量处理所有长字符串,合并为最终的数据框,保留原字符串列。

运行代码后即可得到与期望一致的结果。

内容的提问来源于stack exchange,提问作者CoolGuyHasChillDay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 10:25:22