You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

for-loop结合yield实现星号拼接单词时重复输出后半段的问题排查

问题分析与修复方案

你的核心问题出在join_asterisk函数的遍历逻辑上,以及重复调用该函数导致的重复处理。让我拆解一下问题:

错误原因

  1. 遍历方式导致重复处理:你用zip(words, words[1:], words[2:])每次取三个连续元素,当处理presi, *, dent这一组时,生成了president;但后续遍历到dent, de, la这一组时,dent本身在字典里,会被单独加入结果,这就是它出现在输出里的原因。
  2. 重复调用函数:你分别调用两次join_asterisk(words)生成correct_words和incorrect_words,不仅效率低,还可能引发重复判断的问题。
  3. 未跳过已处理元素:处理完带*的三元组后,没有跳过后续已被拼接的元素,导致它们被单独处理。

修复后的代码

下面是修改后的完整代码,我标注了关键改进点:

from collections import defaultdict
import nltk
from nltk.tokenize import word_tokenize
import re
import os
import sys
from pathlib import Path

def main():
    while True:
        try:
            file_to_open = Path(input("\nPlease, insert your file path: "))
            with open(file_to_open) as f:
                words = word_tokenize(f.read().lower())
            break
        except FileNotFoundError:
            print("\nFile not found. Better try again")
        except IsADirectoryError:
            print("\nIncorrect Directory path.Try again")
    
    word_separator = '*'
    with open('Fr-dictionary2.txt') as fr:
        dic = word_tokenize(fr.read().lower())
    # 转成集合提升查找效率
    dic_set = set(dic)

    def join_asterisk(ary):
        i = 0
        while i < len(ary):
            # 检查是否存在可拼接的三元组
            if i + 2 < len(ary) and ary[i+1] == word_separator:
                combined_word = ary[i] + ary[i+2]
                yield (combined_word, combined_word in dic_set)
                # 跳过已处理的*和后续片段
                i += 3
            else:
                # 处理普通单词,跳过*本身
                current_word = ary[i]
                if current_word != word_separator:
                    yield (current_word, current_word in dic_set)
                i += 1

    # 单次遍历同时收集正确/错误单词
    correct_words = []
    incorrect_words = []
    for word, is_correct in join_asterisk(words):
        if is_correct:
            correct_words.append(word)
        else:
            incorrect_words.append(word)

    text = ' '.join(correct_words)
    print(correct_words)
    print('\n\n', text)
    
    user2 = input('\nWrite text to a file? Type "Y" for yes or "N" for no:')
    if user2 == 'Y':
        text_name = input("name your file.(Ex. 'my_first_file.txt'): ")
        with open(text_name, "w") as out_file:
            out_file.write(text)
    else:
        print('ok')

main()

关键改进点

  1. 索引遍历+跳过已处理元素:改用while循环和索引i,处理完带*的三元组后直接让i +=3,避免*和后续片段被单独处理。
  2. 集合优化查找:把字典转成集合,成员查找效率从O(n)提升到O(1),适合大字典场景。
  3. 单次遍历收集结果:只调用一次join_asterisk,同时填充两个结果列表,避免重复遍历。
  4. 优化文件写入逻辑:仅在用户选择保存时创建并写入文件,避免无效操作。

验证效果

用你的输入示例测试,输出会完全符合期望:

['les', 'engagements', 'du', 'président', 'de', 'la', 'république', 'sont', 'aussi', 'ceux', 'des', 'dirigeants', 'de', 'la', 'société', 'ferroviaire']

内容的提问来源于stack exchange,提问作者Nadia Santos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:37:00