You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中提取文本里最长的正则匹配项?

问题:提取文本中最长的匹配术语(而非第一个匹配项)

我需要基于regex表达式提取文本中的术语,但仅提取最长的匹配项。目前使用str_extract只能提取第一个匹配项(不一定是最长的):

library(stringr)
library(dplyr)
dataframe <- tibble(text = c('how is the biggest ??', 'really amazing stuff'))

dataframe %>% mutate(mymatch = str_extract(text, regex('\\w+')))
# A tibble: 2 x 2
  text                     mymatch
  <chr>                    <chr>  
1 how is the biggest ?? how
2 really amazing stuff     really

尝试使用str_extract_all,但未找到高效语法。期望输出如下:

# A tibble: 2 x 2
  text                     mymatch
  <chr>                    <chr>  
1 how is the biggest ?? biggest
2 really amazing stuff     amazing

解决方案

这个需求很常见,咱们可以基于stringr+dplyr的工具链轻松实现,核心思路是先提取所有匹配项,再在每个文本的匹配结果中筛选出长度最长的那个:

方法一:用str_extract_all + purrr::map_chr

library(dplyr)
library(stringr)
library(purrr)

# 你的原始数据
dataframe <- tibble(text = c('how is the biggest ??', 'really amazing stuff'))

# 处理逻辑
dataframe %>%
  # 第一步:提取当前文本的所有匹配单词,生成列表列
  mutate(all_matches = str_extract_all(text, "\\w+")) %>%
  # 第二步:遍历每个列表,找出长度最长的单词
  mutate(mymatch = map_chr(all_matches, ~ {
    # 计算每个单词的长度,找到最大值的位置,取出对应单词
    .x[which.max(nchar(.x))]
  })) %>%
  # 可选:移除中间生成的all_matches列
  select(-all_matches)

运行后就能得到你想要的结果:

# A tibble: 2 x 2
  text                     mymatch
  <chr>                    <chr>  
1 how is the biggest ?? biggest
2 really amazing stuff     amazing

补充说明

  • 如果同一文本中有多个长度相同的最长单词,这个方法会返回第一个出现的最长单词,符合大多数场景的需求;
  • 如果你不想引入purrr,也可以用基础R的sapply函数替代,写法如下:
dataframe %>%
  mutate(all_matches = str_extract_all(text, "\\w+")) %>%
  mutate(mymatch = sapply(all_matches, function(x) x[which.max(nchar(x))])) %>%
  select(-all_matches)

内容的提问来源于stack exchange,提问作者鈩暿樖樶竼岣结笜

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:27:46