遍历句子匹配多单词时,如何避免重复打印匹配句子?
解决句子匹配单词时的重复输出问题
假设我们有以下句子和待匹配单词:
sentences = ['There are three apples and oranges in the fridge.', 'I forgot the milk.'] wordsMatch = ['apples', 'bananas', 'fridge']
需求是遍历这些句子,只要句子包含任一待匹配单词就输出该句子,但即使句子匹配多个单词也不能重复输出。
比如下面的代码会因为第一个句子同时匹配apples和fridge而重复输出两次:
matchedSentences = [sentence for sentence in sentences for word in wordsMatch if word in sentence] # 输出: ['There are three apples and oranges in the fridge.', 'There are three apples and oranges in the fridge.']
几种可行的解决方案:
1. 使用any()函数的列表推导式(推荐)
这是最简洁高效的方式,对每个句子仅做一次判断,只要存在任一匹配单词就加入结果列表,不会重复:
matchedSentences = [sentence for sentence in sentences if any(word in sentence for word in wordsMatch)] print(matchedSentences)
输出:
['There are three apples and oranges in the fridge.']
any()函数会在找到第一个匹配单词时就停止后续判断,既保证了不重复,又提升了效率。
2. 手动循环+break控制
通过内层循环找到匹配后立即跳出,避免重复添加同一个句子:
matchedSentences = [] for sentence in sentences: for word in wordsMatch: if word in sentence: matchedSentences.append(sentence) break # 找到匹配就终止当前句子的单词检查 print(matchedSentences)
3. 集合去重(不推荐,仅适用于已生成重复列表的场景)
如果已经得到了包含重复项的列表,可以通过转集合再转列表的方式去重,但会打乱原句子的顺序:
matchedSentences = list(set([sentence for sentence in sentences for word in wordsMatch if word in sentence]))
总结
优先选择第一种方法,兼顾可读性和效率;如果需要更灵活的逻辑控制,第二种手动循环的方式更合适;第三种仅作为事后补救的手段,不推荐直接使用。
内容的提问来源于stack exchange,提问作者DarknessPlusPlus
相关产品推荐
相关产品推荐

