You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在文档中高亮含多词元素的关键词集合匹配文本?

多词关键词高效高亮实现方案

正则表达式方案(适合关键词数量不多的场景)

核心思路:优先匹配长关键词,避免短关键词截断长组合的匹配,步骤如下:

  • 将关键词按空格分割后的词数倒序排序,长关键词先匹配
  • 转义关键词中的正则特殊字符,构建忽略大小写的整词匹配模式
  • 用re.sub结合自定义替换逻辑,对匹配内容进行高亮
import re
from termcolor import cprint

def highlight_text(text, keywords):
    # 按关键词词数从多到少排序,优先匹配长组合
    sorted_keywords = sorted(keywords, key=lambda x: len(x.split()), reverse=True)
    # 构建正则模式:整词匹配+忽略大小写
    pattern = re.compile(
        r'\b(' + '|'.join(re.escape(kw) for kw in sorted_keywords) + r')\b',
        re.IGNORECASE
    )
    
    # 替换匹配内容为高亮格式
    def highlight_match(match):
        return f"\033[91m{match.group()}\033[0m"
    
    highlighted_content = pattern.sub(highlight_match, text)
    print(highlighted_content)

# 测试示例
text = "I enjoy Virtual Reality games and virtual experiences"
keywords = ["virtual", "virtual reality"]
highlight_text(text, keywords)

注:如果不需要严格整词匹配(比如允许匹配"virtualized"中的"virtual"),可以去掉正则中的\b边界符。

AC自动机方案(适合大量关键词/大型文档场景)

当关键词数量较多(数百个以上),正则性能会下降,AC自动机可以实现线性时间复杂度的文本扫描,效率更高。借助第三方库ahocorasick快速实现:

首先安装依赖:

pip install ahocorasick

代码实现:

import ahocorasick
from termcolor import cprint

def highlight_with_ahocorasick(text, keywords):
    # 构建AC自动机
    automaton = ahocorasick.Automaton()
    for kw in keywords:
        automaton.add_word(kw.lower(), kw)
    automaton.make_automaton()
    
    # 收集所有匹配的位置与关键词
    matches = []
    text_lower = text.lower()
    for end_idx, kw in automaton.iter(text_lower):
        start_idx = end_idx - len(kw) + 1
        matches.append((start_idx, end_idx, kw))
    
    # 去重处理:保留长且先出现的匹配,移除被覆盖的短匹配
    matches.sort(key=lambda x: (x[0], -(x[1]-x[0]+1)))
    final_matches = []
    last_end_pos = -1
    for start, end, kw in matches:
        if start > last_end_pos:
            final_matches.append((start, end, kw))
            last_end_pos = end
    
    # 拼接高亮后的文本
    result = []
    prev_end = 0
    for start, end, kw in final_matches:
        result.append(text[prev_end:start])
        result.append(f"\033[91m{text[start:end+1]}\033[0m")
        prev_end = end + 1
    result.append(text[prev_end:])
    
    print(''.join(result))

# 测试示例
text = "I enjoy Virtual Reality games and virtual experiences"
keywords = ["virtual", "virtual reality"]
highlight_with_ahocorasick(text, keywords)

内容的提问来源于stack exchange,提问作者almosthavoc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 17:15:53