R语言grepl匹配9万+词汇列表触发内存不足,求低内存方案
低内存实现大词汇表匹配Title的方案
问题背景
我有一个包含16263行数据的表格df:
title date brand big farm house 2022-01-01 A ranch modern 2022-01-01 A town house 2022-01-01 C
另有一个包含94000行数据的表格match_list:
words_for_match farm town clown beach city pink
尝试用正则匹配筛选df中title包含match_list词汇的行时,执行以下代码触发内存不足错误:
match_list <- match_list$words_for_match match_list <- paste(match_list, collapse = "|") match_list <- sprintf("\b(%s)\b", match_list) df %>% filter(grepl(match_list, title))
错误信息:
Problem while computing `..1 = grepl(match_list, subject)`. Caused by error in `grepl()`: ! invalid regular expression, reason 'Out of memory'
仅当match_list缩减至1000行时代码正常运行,需低内存消耗的实现方式。
解决方案
方案1:使用fuzzyjoin包的分词匹配
fuzzyjoin的regex_inner_join可直接基于正则关联两个表,无需拼接超大规模正则表达式,内存占用更可控:
library(fuzzyjoin) library(dplyr) # 给匹配词添加单词边界正则 match_list <- match_list %>% mutate(words_for_match = sprintf("\\b(%s)\\b", words_for_match)) # 执行匹配并去重(避免同一title匹配多个词时重复输出) matched_df <- regex_inner_join(df, match_list, by = c("title" = "words_for_match")) %>% distinct(title, date, brand)
方案2:分块处理匹配词
将94000个匹配词拆分成若干小块,逐块匹配后合并结果,控制单个正则表达式的长度:
library(dplyr) # 按每1000个词分块(可根据内存调整块大小) chunk_size <- 1000 chunks <- split(match_list$words_for_match, ceiling(seq_along(match_list$words_for_match)/chunk_size)) # 逐块匹配并合并结果 matched_df <- lapply(chunks, function(chunk) { pattern <- sprintf("\\b(%s)\\b", paste(chunk, collapse = "|")) df %>% filter(grepl(pattern, title)) }) %>% bind_rows() %>% distinct()
方案3:使用data.table的高效匹配
data.table的%like%运算符内存管理更高效,结合分块进一步降低内存压力:
library(data.table) setDT(df) setDT(match_list) chunk_size <- 1000 chunks <- split(match_list$words_for_match, ceiling(seq_along(match_list$words_for_match)/chunk_size)) matched_df <- rbindlist(lapply(chunks, function(chunk) { pattern <- sprintf("\\b(%s)\\b", paste(chunk, collapse = "|")) df[title %like% pattern] })) %>% unique()
方案4:使用stringi的向量化匹配
stringi的stri_detect_regex支持直接传入匹配词向量,无需拼接成单个正则:
library(stringi) library(dplyr) library(purrr) # 给每个匹配词添加单词边界 patterns <- sprintf("\\b(%s)\\b", match_list$words_for_match) # 批量检查每个title是否存在匹配词 matched_df <- df %>% mutate(has_match = map_lgl(title, ~any(stri_detect_regex(.x, patterns)))) %>% filter(has_match) %>% select(-has_match)
内容的提问来源于stack exchange,提问作者matt_lnrd
相关产品推荐
相关产品推荐

