You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python高效查找长文本中的短语?支持拆分匹配至单词

Python 长文本短语拆分匹配方案

需求可行性分析

这个需求完全可行,核心是从最长到最短遍历所有可能的子短语组合,优先匹配长片段,避免短片段重复匹配。

最优算法思路

采用贪心+子串枚举的思路,步骤如下:

  • 将目标短语按空格拆分为单词列表,比如["there", "is", "a", "group", "of", "students"]。
  • 从当前位置开始,尝试最长的子短语(剩余单词组成的完整片段),检查是否存在于文本中;如果存在,记录该片段并跳过对应单词,继续处理剩余部分。
  • 如果当前长度的子短语不存在,就缩短长度(减少末尾一个单词),重复检查;直到子短语长度为1(单个单词),若存在则记录。
  • 重复上述过程,直到处理完所有目标单词。

这种思路保证了优先匹配最长有效子短语,避免拆分过度,同时减少不必要的重复检查。

代码实现示例

def find_max_matches(text, target_phrase):
    target_words = target_phrase.split()
    matches = []
    total_words = len(target_words)
    current_idx = 0
    
    while current_idx < total_words:
        found = False
        # 从当前位置开始,尝试最长到最短的子短语
        for phrase_length in range(total_words - current_idx, 0, -1):
            current_phrase = ' '.join(target_words[current_idx:current_idx+phrase_length])
            if current_phrase in text:
                matches.append(current_phrase)
                current_idx += phrase_length
                found = True
                break
        if not found:
            # 单个单词也未匹配,跳过该单词(可按需记录未匹配项)
            current_idx += 1
    return matches

# 测试示例
sample_text = "there are some people, is a group of teachers and students playing outside"
target_phrase = "there is a group of students"
print(find_max_matches(sample_text, target_phrase))
# 输出:['there', 'is a group of', 'students']

效率优化建议

  • 对于超长篇文本,直接用in操作会重复扫描文本,建议预处理为单词列表,用滑动窗口比对单词序列,避免字符串拼接和重复扫描:
    def find_max_matches_optimized(text, target_phrase):
        text_words = text.split()
        target_words = target_phrase.split()
        matches = []
        target_len = len(target_words)
        current_idx = 0
        
        while current_idx < target_len:
            found = False
            for match_len in range(target_len - current_idx, 0, -1):
                # 在文本单词列表中滑动查找匹配的单词序列
                for window_start in range(len(text_words) - match_len + 1):
                    if text_words[window_start:window_start+match_len] == target_words[current_idx:current_idx+match_len]:
                        matches.append(' '.join(target_words[current_idx:current_idx+match_len]))
                        current_idx += match_len
                        found = True
                        break
                if found:
                    break
            if not found:
                current_idx += 1
        return matches
    
  • 若需处理大规模文本匹配,可结合KMP算法或Aho-Corasick自动机预处理目标短语的所有可能子序列,进一步减少比对次数。

内容的提问来源于stack exchange,提问作者Val

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 10:32:36