You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Python的filter返回含匹配关键词与句子的元组列表?

问题:如何让关键词匹配函数返回(关键词,句子)元组列表?

你手头有两个超大列表——150万条句子的list_of_sents和长度相近的关键词列表list_of_words,需要高效找出包含关键词的句子,而且极度在意计算效率,想尽量避免冗余的for循环。

先看你的示例数据:

list_of_words = ['Turin', 'Milan'] 
list_of_sents = ['This is a sent about turin.', 'This is a sent about manufacturing.']

你已经写了一个高效的匹配函数,但它只能返回匹配的句子列表:

import re

def find_keyword_comments(test_comments,test_keywords): 
    keywords = '|'.join(test_keywords) 
    word = re.compile(r"^.*\b({})\b.*$".format(keywords), re.I) 
    newlist = filter(word.match, test_comments) 
    final = list(newlist) 
    return final

现有输出:

['This is a sent about turin.']

而你想要的是包含原关键词和对应句子的元组列表,比如:

[('Turin', 'This is a sent about turin.')]

解决方案:捕获匹配关键词+快速映射还原

核心思路是:让正则捕获匹配到的关键词,同时用一个字典快速把匹配到的内容(可能是小写)映射回原关键词的大小写,全程只遍历一次句子列表,保证效率。

优化后的函数(Python3.8+,用海象运算符简化)

import re

def find_keyword_comments(test_comments, test_keywords):
    # 建立小写关键词到原关键词的映射,O(1)查找还原大小写
    keyword_lower_map = {kw.lower(): kw for kw in test_keywords}
    # 转义关键词中的特殊正则字符,避免语法错误
    keywords_pattern = '|'.join(re.escape(kw) for kw in test_keywords)
    # 预编译正则,只做一次编译,提升效率
    regex = re.compile(r"\b({})\b".format(keywords_pattern), re.I)
    
    # 列表推导式遍历句子,匹配到就返回(原关键词,句子)元组
    return [
        (keyword_lower_map[match.group(1).lower()], sent)
        for sent in test_comments
        if (match := regex.search(sent))  # 海象运算符,一次匹配判断+捕获结果
    ]

兼容低版本Python的写法(无海象运算符)

import re

def find_keyword_comments(test_comments, test_keywords):
    keyword_lower_map = {kw.lower(): kw for kw in test_keywords}
    keywords_pattern = '|'.join(re.escape(kw) for kw in test_keywords)
    regex = re.compile(r"\b({})\b".format(keywords_pattern), re.I)
    
    final_list = []
    for sent in test_comments:
        match_result = regex.search(sent)
        if match_result:
            # 还原成原关键词的大小写
            matched_keyword = keyword_lower_map[match_result.group(1).lower()]
            final_list.append((matched_keyword, sent))
    return final_list

为什么这个方案高效?

  • 字典映射:keyword_lower_map让我们可以O(1)时间还原原关键词的大小写,避免循环遍历关键词列表去匹配。
  • 预编译正则:只编译一次正则表达式,避免重复编译的开销。
  • 单次遍历:只遍历一次句子列表,每个句子的正则search操作是高效的(比你原来的match更合适,因为match是从句子开头匹配,search会找任意位置的关键词)。
  • 转义处理:用re.escape处理关键词,避免关键词里的特殊字符(比如.、*)破坏正则语法。

测试你的示例数据,这个函数会返回你想要的结果:

[('Turin', 'This is a sent about turin.')]

内容的提问来源于stack exchange,提问作者JRR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:43:21