如何让str_which按最匹配优先级返回字符串匹配结果?
解决str_which无法按最长匹配优先返回结果的问题
核心思路是:先筛选出所有能匹配目标字符串的pattern,再按匹配字符串的长度从长到短排序,最后返回对应的原始pattern索引。
自定义实现函数
str_which_longest_first <- function(string, pattern) { # 获取所有匹配的pattern索引 match_idx <- str_which(string, pattern) # 提取匹配到的pattern内容 matched_patterns <- pattern[match_idx] # 按字符串长度降序排序,得到排序后的原始索引 sorted_idx <- match_idx[order(nchar(matched_patterns), decreasing = TRUE)] sorted_idx }
验证示例
用你提供的测试数据验证效果:
df <- data.frame(city = c("beautiful city", "cool beautiful city", "strange local", "noname local")) pattern <- c("cool beautiful city", "name local", "noname local", "beautiful city") # 测试第一个城市 str_which_longest_first(df$city[1], pattern) # 输出: [1] 4 1 # 测试第四个城市 str_which_longest_first(df$city[4], pattern) # 输出: [1] 3 2
原理说明
str_which(string, pattern):先完成基础的匹配筛选,得到所有符合条件的pattern索引order(nchar(matched_patterns), decreasing = TRUE):对匹配到的pattern按字符长度降序排序,确保更长、更精确的匹配结果排在最前面- 最终返回排序后的原始pattern索引,完全满足你要的“最匹配优先”需求
内容的提问来源于stack exchange,提问作者alexgatsby
相关产品推荐
相关产品推荐

