如何在R语言中提取文本里最长的正则匹配项?
问题:提取文本中最长的匹配术语(而非第一个匹配项)
我需要基于regex表达式提取文本中的术语,但仅提取最长的匹配项。目前使用str_extract只能提取第一个匹配项(不一定是最长的):
library(stringr) library(dplyr) dataframe <- tibble(text = c('how is the biggest ??', 'really amazing stuff')) dataframe %>% mutate(mymatch = str_extract(text, regex('\\w+'))) # A tibble: 2 x 2 text mymatch <chr> <chr> 1 how is the biggest ?? how 2 really amazing stuff really
尝试使用str_extract_all,但未找到高效语法。期望输出如下:
# A tibble: 2 x 2 text mymatch <chr> <chr> 1 how is the biggest ?? biggest 2 really amazing stuff amazing
解决方案
这个需求很常见,咱们可以基于stringr+dplyr的工具链轻松实现,核心思路是先提取所有匹配项,再在每个文本的匹配结果中筛选出长度最长的那个:
方法一:用str_extract_all + purrr::map_chr
library(dplyr) library(stringr) library(purrr) # 你的原始数据 dataframe <- tibble(text = c('how is the biggest ??', 'really amazing stuff')) # 处理逻辑 dataframe %>% # 第一步:提取当前文本的所有匹配单词,生成列表列 mutate(all_matches = str_extract_all(text, "\\w+")) %>% # 第二步:遍历每个列表,找出长度最长的单词 mutate(mymatch = map_chr(all_matches, ~ { # 计算每个单词的长度,找到最大值的位置,取出对应单词 .x[which.max(nchar(.x))] })) %>% # 可选:移除中间生成的all_matches列 select(-all_matches)
运行后就能得到你想要的结果:
# A tibble: 2 x 2 text mymatch <chr> <chr> 1 how is the biggest ?? biggest 2 really amazing stuff amazing
补充说明
- 如果同一文本中有多个长度相同的最长单词,这个方法会返回第一个出现的最长单词,符合大多数场景的需求;
- 如果你不想引入
purrr,也可以用基础R的sapply函数替代,写法如下:
dataframe %>% mutate(all_matches = str_extract_all(text, "\\w+")) %>% mutate(mymatch = sapply(all_matches, function(x) x[which.max(nchar(x))])) %>% select(-all_matches)
内容的提问来源于stack exchange,提问作者鈩暿樖樶竼岣结笜
相关产品推荐
相关产品推荐

