You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

开发函数筛选含目标单词且无邻接字母数字的百万级句子

嘿,针对你这个百万级句子和单词列表的匹配需求,我整理了几个兼顾准确性和性能的实现方案,应该能完美解决你的问题:

核心需求拆解

首先明确你的核心要求:要找出包含目标单词,且单词前后没有字母数字字符的句子,而且数据量高达150万级,所以实现必须兼顾精准度和执行效率。

比如示例里的"turin"能匹配目标词"Turin",但如果是"Turin123"或者"xTurin"就不符合要求——因为前后有字母/数字。

高效实现方案(基础版)

这个版本用正则预编译+批量匹配的思路,把时间复杂度从O(n*m)降到O(n),完全适配百万级数据:

import re

def filter_target_sentences(list_of_sents, list_of_words):
    # 处理空输入的边界情况
    if not list_of_sents or not list_of_words:
        return []
    
    # 转义每个目标单词,避免正则特殊字符(比如".", "*")导致的解析错误
    escaped_words = [re.escape(word) for word in list_of_words]
    # 构建正则模式:用负向断言确保单词前后无字母数字,不区分大小写
    pattern = re.compile(
        r'(?<!\w)({})(?!\w)'.format('|'.join(escaped_words)),
        re.IGNORECASE
    )
    
    # 用列表推导式快速筛选,底层是C实现,比普通循环快很多
    return [sentence for sentence in list_of_sents if pattern.search(sentence)]

测试示例

list_of_words = ['Turin', 'Milan']
list_of_sents = [
    'This is a sent about turin.',
    'This is a sent about manufacturing.',
    'I love Milan!',
    'Turin123 is invalid',
    'xMilanx is also invalid'
]

print(filter_target_sentences(list_of_sents, list_of_words))
# 输出:['This is a sent about turin.', 'I love Milan!']
性能优化进阶(并行版)

如果你的机器是多核CPU,处理150万句子可以用多进程并行提速,把任务拆分到多个核心同时处理:

from multiprocessing import Pool
import re

# 初始化子进程的全局变量,传递预编译好的正则
def init_worker(compiled_pattern):
    global pattern
    pattern = compiled_pattern

# 单个句子的匹配函数
def is_match(sentence):
    return pattern.search(sentence) is not None

def filter_sentences_parallel(list_of_sents, list_of_words, num_workers=4):
    if not list_of_sents or not list_of_words:
        return []
    
    escaped_words = [re.escape(word) for word in list_of_words]
    compiled_pattern = re.compile(
        r'(?<!\w)({})(?!\w)'.format('|'.join(escaped_words)),
        re.IGNORECASE
    )
    
    # 启动多进程池,并行检查所有句子
    with Pool(num_workers, initializer=init_worker, initargs=(compiled_pattern,)) as pool:
        match_results = pool.map(is_match, list_of_sents)
    
    # 筛选出匹配的句子
    return [sent for sent, matched in zip(list_of_sents, match_results) if matched]
关键细节说明
  • 为什么用(?<!\w)和(?!\w)?
    \w等价于[a-zA-Z0-9_],负向断言(?<!\w)确保目标单词的前面不是字母/数字/下划线,(?!\w)确保后面也不是——完美符合你“单词前后无字母数字”的要求,比\b更精准(\b会把下划线当成单词边界,不符合你的需求)。
  • 正则预编译的重要性:re.compile只执行一次,避免每个句子匹配都重复编译正则,节省大量时间。
  • 单词转义:如果目标单词包含正则特殊字符(比如Mr.Smith里的.),re.escape会自动转义,避免正则解析出错。
  • 内存优化:如果句子列表大到内存吃不消,可以把列表推导式改成生成器表达式:(sent for sent in list_of_sents if pattern.search(sent)),这样不会一次性加载所有结果到内存。

内容的提问来源于stack exchange,提问作者JRR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:41:34