for-loop结合yield实现星号拼接单词时重复输出后半段的问题排查
问题分析与修复方案
你的核心问题出在join_asterisk函数的遍历逻辑上,以及重复调用该函数导致的重复处理。让我拆解一下问题:
错误原因
- 遍历方式导致重复处理:你用
zip(words, words[1:], words[2:])每次取三个连续元素,当处理presi,*,dent这一组时,生成了president;但后续遍历到dent,de,la这一组时,dent本身在字典里,会被单独加入结果,这就是它出现在输出里的原因。 - 重复调用函数:你分别调用两次
join_asterisk(words)生成correct_words和incorrect_words,不仅效率低,还可能引发重复判断的问题。 - 未跳过已处理元素:处理完带
*的三元组后,没有跳过后续已被拼接的元素,导致它们被单独处理。
修复后的代码
下面是修改后的完整代码,我标注了关键改进点:
from collections import defaultdict import nltk from nltk.tokenize import word_tokenize import re import os import sys from pathlib import Path def main(): while True: try: file_to_open = Path(input("\nPlease, insert your file path: ")) with open(file_to_open) as f: words = word_tokenize(f.read().lower()) break except FileNotFoundError: print("\nFile not found. Better try again") except IsADirectoryError: print("\nIncorrect Directory path.Try again") word_separator = '*' with open('Fr-dictionary2.txt') as fr: dic = word_tokenize(fr.read().lower()) # 转成集合提升查找效率 dic_set = set(dic) def join_asterisk(ary): i = 0 while i < len(ary): # 检查是否存在可拼接的三元组 if i + 2 < len(ary) and ary[i+1] == word_separator: combined_word = ary[i] + ary[i+2] yield (combined_word, combined_word in dic_set) # 跳过已处理的*和后续片段 i += 3 else: # 处理普通单词,跳过*本身 current_word = ary[i] if current_word != word_separator: yield (current_word, current_word in dic_set) i += 1 # 单次遍历同时收集正确/错误单词 correct_words = [] incorrect_words = [] for word, is_correct in join_asterisk(words): if is_correct: correct_words.append(word) else: incorrect_words.append(word) text = ' '.join(correct_words) print(correct_words) print('\n\n', text) user2 = input('\nWrite text to a file? Type "Y" for yes or "N" for no:') if user2 == 'Y': text_name = input("name your file.(Ex. 'my_first_file.txt'): ") with open(text_name, "w") as out_file: out_file.write(text) else: print('ok') main()
关键改进点
- 索引遍历+跳过已处理元素:改用
while循环和索引i,处理完带*的三元组后直接让i +=3,避免*和后续片段被单独处理。 - 集合优化查找:把字典转成集合,成员查找效率从O(n)提升到O(1),适合大字典场景。
- 单次遍历收集结果:只调用一次
join_asterisk,同时填充两个结果列表,避免重复遍历。 - 优化文件写入逻辑:仅在用户选择保存时创建并写入文件,避免无效操作。
验证效果
用你的输入示例测试,输出会完全符合期望:
['les', 'engagements', 'du', 'président', 'de', 'la', 'république', 'sont', 'aussi', 'ceux', 'des', 'dirigeants', 'de', 'la', 'société', 'ferroviaire']
内容的提问来源于stack exchange,提问作者Nadia Santos
相关产品推荐
相关产品推荐

