You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中多行字符串匹配逻辑问题求助及代码优化需求

Python多行字符串匹配逻辑修正问题

需求是将search_sentence中的单词依次匹配到raw_sentences的每一行句子中,现有代码运行结果不符合预期,需要调整代码实现期望输出。

现有代码

search_sentence = "for learning distinctive features among"

raw_sentences = ["methods it is possible to among recognize Instagram filters and at-", # 01
                 "tenuate the sensor pattern noise signal in images. Amerini", # 02
                 "et al. [10] introduced a CNN for learning distinctive features", # 03
                 "among social networks. for learning distinctive features among from the histogram of the discrete co-", # 04
                 "sine transform (DCT) coefficients and the noise residual of", # 05
                 "the images. Phan et al. [11] proposed a method to track mul-", # 06
                 "tiple image sharing on social networks by using a CNN for ar-", # 07
                 "chitecture able to learn", # 08
                 "et al. [10] introduced a CNN for learning distinctive features among it is possible to among recognize Instagram filters", # 09
                 "and at- tenuate xx"] # 10


def longest_intersection(string1, string2):
    list1 = string1.split()
    list2 = string2.split()
    intersection = []
    for word in list1:
        if word in list2 and word == list2[0]:
            intersection.append(word)
            list2.remove(word)
    if " ".join(intersection) in search_sentence:
        return intersection


for line in raw_sentences:
    one_line_match = ' '.join(longest_intersection(line.strip(), search_sentence))

    if one_line_match != "" and one_line_match[0] == search_sentence[0]:
        print(one_line_match)
        search_sentence = search_sentence.replace(one_line_match, "").strip()
        if search_sentence == "":
            search_sentence = "for learning distinctive features among"
    else:
        print("[no matched sentences!]")
        search_sentence = "for learning distinctive features among"

当前输出

[no matched sentences!]
[no matched sentences!]
for learning distinctive features
among
[no matched sentences!]
[no matched sentences!]
for
[no matched sentences!]
for learning distinctive features among
[no matched sentences!]

期望输出

[no matched sentences!]
[no matched sentences!]
for learning distinctive features
among for learning distinctive features among
[no matched sentences!]
[no matched sentences!]
[no matched sentences!]
[no matched sentences!]
for learning distinctive features among
[no matched sentences!]

问题分析与修正代码

原代码存在两个核心问题:

  1. 仅匹配search_sentence开头的连续单词,无法处理同一行内多次出现完整匹配序列的场景
  2. 匹配后重置搜索序列的逻辑僵化,未考虑一行内可完成剩余匹配+多次完整匹配的情况

修正后的代码:

search_sentence = "for learning distinctive features among"
target_words = search_sentence.split()
target_len = len(target_words)

raw_sentences = ["methods it is possible to among recognize Instagram filters and at-", # 01
                 "tenuate the sensor pattern noise signal in images. Amerini", # 02
                 "et al. [10] introduced a CNN for learning distinctive features", # 03
                 "among social networks. for learning distinctive features among from the histogram of the discrete co-", # 04
                 "sine transform (DCT) coefficients and the noise residual of", # 05
                 "the images. Phan et al. [11] proposed a method to track mul-", # 06
                 "tiple image sharing on social networks by using a CNN for ar-", # 07
                 "chitecture able to learn", # 08
                 "et al. [10] introduced a CNN for learning distinctive features among it is possible to among recognize Instagram filters", # 09
                 "and at- tenuate xx"] # 10


def process_line(line_words, current_search, target_seq):
    matches = []
    # 优先处理剩余的待匹配序列
    if current_search:
        prefix = []
        for word in line_words:
            if prefix and len(prefix) < len(current_search) and word == current_search[len(prefix)]:
                prefix.append(word)
            elif not prefix and word == current_search[0]:
                prefix.append(word)
            else:
                break
        if prefix:
            matches.append(' '.join(prefix))
            # 更新剩余待匹配序列
            current_search = current_search[len(prefix):]
            if not current_search:
                current_search = target_seq.copy()
    
    # 查找当前行中所有完整的目标序列匹配
    seq_len = len(target_seq)
    for i in range(len(line_words) - seq_len + 1):
        if line_words[i:i+seq_len] == target_seq:
            matches.append(' '.join(target_seq))
            # 跳过已匹配的单词,避免重叠匹配
            i += seq_len - 1
    
    return matches, current_search


current_search = target_words.copy()

for line in raw_sentences:
    line_words = line.strip().split()
    line_matches, current_search = process_line(line_words, current_search, target_words)
    
    if line_matches:
        print(' '.join(line_matches))
    else:
        print("[no matched sentences!]")
        # 无匹配时重置待匹配序列
        current_search = target_words.copy()

修正说明

  1. 将目标序列拆分为单词列表,便于逐段匹配和完整序列比对
  2. 维护current_search变量跟踪剩余待匹配的单词,匹配完成后自动重置为完整目标序列
  3. 处理每行时先完成剩余匹配,再查找所有完整序列匹配,满足第4行的混合匹配需求
  4. 跳过已匹配的单词段,避免同一行内的重叠匹配问题

内容的提问来源于stack exchange,提问作者CAO RUI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 12:45:36