Python中多行字符串匹配逻辑问题求助及代码优化需求
Python多行字符串匹配逻辑修正问题
需求是将search_sentence中的单词依次匹配到raw_sentences的每一行句子中,现有代码运行结果不符合预期,需要调整代码实现期望输出。
现有代码
search_sentence = "for learning distinctive features among" raw_sentences = ["methods it is possible to among recognize Instagram filters and at-", # 01 "tenuate the sensor pattern noise signal in images. Amerini", # 02 "et al. [10] introduced a CNN for learning distinctive features", # 03 "among social networks. for learning distinctive features among from the histogram of the discrete co-", # 04 "sine transform (DCT) coefficients and the noise residual of", # 05 "the images. Phan et al. [11] proposed a method to track mul-", # 06 "tiple image sharing on social networks by using a CNN for ar-", # 07 "chitecture able to learn", # 08 "et al. [10] introduced a CNN for learning distinctive features among it is possible to among recognize Instagram filters", # 09 "and at- tenuate xx"] # 10 def longest_intersection(string1, string2): list1 = string1.split() list2 = string2.split() intersection = [] for word in list1: if word in list2 and word == list2[0]: intersection.append(word) list2.remove(word) if " ".join(intersection) in search_sentence: return intersection for line in raw_sentences: one_line_match = ' '.join(longest_intersection(line.strip(), search_sentence)) if one_line_match != "" and one_line_match[0] == search_sentence[0]: print(one_line_match) search_sentence = search_sentence.replace(one_line_match, "").strip() if search_sentence == "": search_sentence = "for learning distinctive features among" else: print("[no matched sentences!]") search_sentence = "for learning distinctive features among"
当前输出
[no matched sentences!] [no matched sentences!] for learning distinctive features among [no matched sentences!] [no matched sentences!] for [no matched sentences!] for learning distinctive features among [no matched sentences!]
期望输出
[no matched sentences!] [no matched sentences!] for learning distinctive features among for learning distinctive features among [no matched sentences!] [no matched sentences!] [no matched sentences!] [no matched sentences!] for learning distinctive features among [no matched sentences!]
问题分析与修正代码
原代码存在两个核心问题:
- 仅匹配
search_sentence开头的连续单词,无法处理同一行内多次出现完整匹配序列的场景 - 匹配后重置搜索序列的逻辑僵化,未考虑一行内可完成剩余匹配+多次完整匹配的情况
修正后的代码:
search_sentence = "for learning distinctive features among" target_words = search_sentence.split() target_len = len(target_words) raw_sentences = ["methods it is possible to among recognize Instagram filters and at-", # 01 "tenuate the sensor pattern noise signal in images. Amerini", # 02 "et al. [10] introduced a CNN for learning distinctive features", # 03 "among social networks. for learning distinctive features among from the histogram of the discrete co-", # 04 "sine transform (DCT) coefficients and the noise residual of", # 05 "the images. Phan et al. [11] proposed a method to track mul-", # 06 "tiple image sharing on social networks by using a CNN for ar-", # 07 "chitecture able to learn", # 08 "et al. [10] introduced a CNN for learning distinctive features among it is possible to among recognize Instagram filters", # 09 "and at- tenuate xx"] # 10 def process_line(line_words, current_search, target_seq): matches = [] # 优先处理剩余的待匹配序列 if current_search: prefix = [] for word in line_words: if prefix and len(prefix) < len(current_search) and word == current_search[len(prefix)]: prefix.append(word) elif not prefix and word == current_search[0]: prefix.append(word) else: break if prefix: matches.append(' '.join(prefix)) # 更新剩余待匹配序列 current_search = current_search[len(prefix):] if not current_search: current_search = target_seq.copy() # 查找当前行中所有完整的目标序列匹配 seq_len = len(target_seq) for i in range(len(line_words) - seq_len + 1): if line_words[i:i+seq_len] == target_seq: matches.append(' '.join(target_seq)) # 跳过已匹配的单词,避免重叠匹配 i += seq_len - 1 return matches, current_search current_search = target_words.copy() for line in raw_sentences: line_words = line.strip().split() line_matches, current_search = process_line(line_words, current_search, target_words) if line_matches: print(' '.join(line_matches)) else: print("[no matched sentences!]") # 无匹配时重置待匹配序列 current_search = target_words.copy()
修正说明
- 将目标序列拆分为单词列表,便于逐段匹配和完整序列比对
- 维护
current_search变量跟踪剩余待匹配的单词,匹配完成后自动重置为完整目标序列 - 处理每行时先完成剩余匹配,再查找所有完整序列匹配,满足第4行的混合匹配需求
- 跳过已匹配的单词段,避免同一行内的重叠匹配问题
内容的提问来源于stack exchange,提问作者CAO RUI
相关产品推荐
相关产品推荐

