You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言grepl匹配9万+词汇列表触发内存不足,求低内存方案

低内存实现大词汇表匹配Title的方案

问题背景

我有一个包含16263行数据的表格df:

title             date            brand
big farm house    2022-01-01      A
ranch modern      2022-01-01      A
town house        2022-01-01      C

另有一个包含94000行数据的表格match_list:

words_for_match
farm
town
clown
beach
city
pink

尝试用正则匹配筛选df中title包含match_list词汇的行时,执行以下代码触发内存不足错误:

match_list <- match_list$words_for_match
match_list <- paste(match_list, collapse = "|")
match_list <- sprintf("\b(%s)\b", match_list)

df %>% 
  filter(grepl(match_list, title))

错误信息:

Problem while computing `..1 = grepl(match_list, subject)`.
Caused by error in `grepl()`:
! invalid regular expression, reason 'Out of memory'

仅当match_list缩减至1000行时代码正常运行,需低内存消耗的实现方式。


解决方案

方案1:使用fuzzyjoin包的分词匹配

fuzzyjoin的regex_inner_join可直接基于正则关联两个表,无需拼接超大规模正则表达式,内存占用更可控:

library(fuzzyjoin)
library(dplyr)

# 给匹配词添加单词边界正则
match_list <- match_list %>%
  mutate(words_for_match = sprintf("\\b(%s)\\b", words_for_match))

# 执行匹配并去重(避免同一title匹配多个词时重复输出)
matched_df <- regex_inner_join(df, match_list, by = c("title" = "words_for_match")) %>%
  distinct(title, date, brand)

方案2:分块处理匹配词

将94000个匹配词拆分成若干小块,逐块匹配后合并结果,控制单个正则表达式的长度:

library(dplyr)

# 按每1000个词分块(可根据内存调整块大小)
chunk_size <- 1000
chunks <- split(match_list$words_for_match, ceiling(seq_along(match_list$words_for_match)/chunk_size))

# 逐块匹配并合并结果
matched_df <- lapply(chunks, function(chunk) {
  pattern <- sprintf("\\b(%s)\\b", paste(chunk, collapse = "|"))
  df %>% filter(grepl(pattern, title))
}) %>% bind_rows() %>% distinct()

方案3:使用data.table的高效匹配

data.table的%like%运算符内存管理更高效,结合分块进一步降低内存压力:

library(data.table)

setDT(df)
setDT(match_list)

chunk_size <- 1000
chunks <- split(match_list$words_for_match, ceiling(seq_along(match_list$words_for_match)/chunk_size))

matched_df <- rbindlist(lapply(chunks, function(chunk) {
  pattern <- sprintf("\\b(%s)\\b", paste(chunk, collapse = "|"))
  df[title %like% pattern]
})) %>% unique()

方案4:使用stringi的向量化匹配

stringi的stri_detect_regex支持直接传入匹配词向量,无需拼接成单个正则:

library(stringi)
library(dplyr)
library(purrr)

# 给每个匹配词添加单词边界
patterns <- sprintf("\\b(%s)\\b", match_list$words_for_match)

# 批量检查每个title是否存在匹配词
matched_df <- df %>%
  mutate(has_match = map_lgl(title, ~any(stri_detect_regex(.x, patterns)))) %>%
  filter(has_match) %>%
  select(-has_match)

内容的提问来源于stack exchange,提问作者matt_lnrd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 02:40:40