开发函数筛选含目标单词且无邻接字母数字的百万级句子
嘿,针对你这个百万级句子和单词列表的匹配需求,我整理了几个兼顾准确性和性能的实现方案,应该能完美解决你的问题:
核心需求拆解
首先明确你的核心要求:要找出包含目标单词,且单词前后没有字母数字字符的句子,而且数据量高达150万级,所以实现必须兼顾精准度和执行效率。
比如示例里的"turin"能匹配目标词"Turin",但如果是"Turin123"或者"xTurin"就不符合要求——因为前后有字母/数字。
高效实现方案(基础版)
这个版本用正则预编译+批量匹配的思路,把时间复杂度从O(n*m)降到O(n),完全适配百万级数据:
import re def filter_target_sentences(list_of_sents, list_of_words): # 处理空输入的边界情况 if not list_of_sents or not list_of_words: return [] # 转义每个目标单词,避免正则特殊字符(比如".", "*")导致的解析错误 escaped_words = [re.escape(word) for word in list_of_words] # 构建正则模式:用负向断言确保单词前后无字母数字,不区分大小写 pattern = re.compile( r'(?<!\w)({})(?!\w)'.format('|'.join(escaped_words)), re.IGNORECASE ) # 用列表推导式快速筛选,底层是C实现,比普通循环快很多 return [sentence for sentence in list_of_sents if pattern.search(sentence)]
测试示例
list_of_words = ['Turin', 'Milan'] list_of_sents = [ 'This is a sent about turin.', 'This is a sent about manufacturing.', 'I love Milan!', 'Turin123 is invalid', 'xMilanx is also invalid' ] print(filter_target_sentences(list_of_sents, list_of_words)) # 输出:['This is a sent about turin.', 'I love Milan!']
性能优化进阶(并行版)
如果你的机器是多核CPU,处理150万句子可以用多进程并行提速,把任务拆分到多个核心同时处理:
from multiprocessing import Pool import re # 初始化子进程的全局变量,传递预编译好的正则 def init_worker(compiled_pattern): global pattern pattern = compiled_pattern # 单个句子的匹配函数 def is_match(sentence): return pattern.search(sentence) is not None def filter_sentences_parallel(list_of_sents, list_of_words, num_workers=4): if not list_of_sents or not list_of_words: return [] escaped_words = [re.escape(word) for word in list_of_words] compiled_pattern = re.compile( r'(?<!\w)({})(?!\w)'.format('|'.join(escaped_words)), re.IGNORECASE ) # 启动多进程池,并行检查所有句子 with Pool(num_workers, initializer=init_worker, initargs=(compiled_pattern,)) as pool: match_results = pool.map(is_match, list_of_sents) # 筛选出匹配的句子 return [sent for sent, matched in zip(list_of_sents, match_results) if matched]
关键细节说明
- 为什么用
(?<!\w)和(?!\w)?\w等价于[a-zA-Z0-9_],负向断言(?<!\w)确保目标单词的前面不是字母/数字/下划线,(?!\w)确保后面也不是——完美符合你“单词前后无字母数字”的要求,比\b更精准(\b会把下划线当成单词边界,不符合你的需求)。 - 正则预编译的重要性:
re.compile只执行一次,避免每个句子匹配都重复编译正则,节省大量时间。 - 单词转义:如果目标单词包含正则特殊字符(比如
Mr.Smith里的.),re.escape会自动转义,避免正则解析出错。 - 内存优化:如果句子列表大到内存吃不消,可以把列表推导式改成生成器表达式:
(sent for sent in list_of_sents if pattern.search(sent)),这样不会一次性加载所有结果到内存。
内容的提问来源于stack exchange,提问作者JRR
相关产品推荐
相关产品推荐

