You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

保加利亚语文本生成脚本陷入无限循环问题求助

保加利亚语文本生成器无限循环问题排查

问题描述

开发保加利亚语简易文本生成器时,代码陷入无限循环。调试发现生成文本的嵌套for循环中print语句从未执行,无法向text列表添加新单词。原代码使用reuters.sents()可正常运行,现基于自定义分词和训练文本出现异常。

相关代码

主代码

from tokenization import tokenize_bulgarian_text
from nltk import bigrams, trigrams
from collections import Counter, defaultdict
import random

with open('IvanVazov1.txt', 'r', encoding='utf-8') as f:
    data = f.read()

# Tokenize text
tokenized_sentences = tokenize_bulgarian_text(data)
print(tokenized_sentences)

# Create a placeholder for model
model = defaultdict(lambda: defaultdict(lambda: 0))

# Count frequency of co-occurance  
for sentence in tokenized_sentences:
    for w1, w2, w3 in trigrams(sentence, pad_right=True, pad_left=True):
        print("Trigram:", (w1, w2, w3))
        model[(w1, w2)][w3] += 1
 
# Transform the counts to probabilities
for w1_w2 in model:
    total_count = float(sum(model[w1_w2].values()))
    for w3 in model[w1_w2]:
        model[w1_w2][w3] /= total_count
print(model)


# starting words
text = ["беше"]
sentence_finished = False
 
print("Starting words:", text)

while not sentence_finished:
    # select a random probability threshold  
    r = random.random()
    accumulator = .0
    for word in model[tuple(text[-2:])].keys():
        # print statement never executes
        print(f"word: {word}, probability: {model[tuple(text[-2:])][word]}, accumulator: {accumulator}")
        accumulator += model[tuple(text[-2:])][word]
        # select words that are above the probability threshold
        if accumulator >= r:
            text.append(word)
            break

    if text[-2:] == [None, None]:
        print("End of sentence.")
        sentence_finished = True
 
print (' '.join([t for t in text if t]))

tokenization.py代码

import nltk
from nltk.tokenize import sent_tokenize, word_tokenize
import string
from spacy.lang.bg.stop_words import STOP_WORDS as bg_stopwords


def tokenize_bulgarian_text(text):
    extra_stopwords = {"—", "“", "„", "не", "та", "па"}
    bg_stopwords.update(extra_stopwords)
    # Remove punctuation and lowercase all letters
    text = text.translate(str.maketrans('', '', string.punctuation))
    text = text.lower()
    
    # Split the text into individual sentences
    sentences = sent_tokenize(text)

    # Tokenize and filter each sentence
    filtered_tokens = []
    for sentence in sentences:
        # Split the sentence into individual words or tokens
        tokens = word_tokenize(sentence)
        # Remove any stop words from the tokenized text
        filtered_sentence = [word for word in tokens if word not in bg_stopwords]
        filtered_tokens.append(filtered_sentence)

    return filtered_tokens

训练文本

Нощта беше влажна и мрачна и браилските улици пустееха. Студената декемврийска мъгла, която обикновено пада покрай бреговете на Дунава, се беше напластила в една от главните улици на града и задушаваше с отровния си дъх последните минувачи, които бързаха да се приберат у дома си. 

问题分析与解决方案

1. 核心问题:初始查询键不存在

你的text初始值只有单个单词["беше"],执行model[tuple(text[-2:])]时,得到的是单元素元组("беше",),但你的trigram模型中所有键都是二元组(比如(None, None)、(None, "нощта")、("нощта", "беше")等),因此这个查询会返回一个空的defaultdict,导致嵌套for循环根本不会执行,text永远无法新增单词,最终陷入无限循环。

解决方法:
将初始text设置为训练数据中存在的连续二元组,比如从训练文本中选取["нощта", "беше"],这样text[-2:]就能匹配模型中存在的键("нощта", "беше"),从而正常获取后续单词:

# starting words
text = ["нощта", "беше"]

2. 次要问题:分词步骤顺序错误

当前分词函数先全局移除所有标点,再执行sent_tokenize分割句子。但sent_tokenize依赖标点(如句号)识别句子边界,移除标点后会导致句子分割错误,进而影响trigram模型的构建。

解决方法:调整分词步骤顺序,先分割句子,再对每个句子单独处理标点和小写:

def tokenize_bulgarian_text(text):
    extra_stopwords = {"—", "“", "„", "не", "та", "па"}
    bg_stopwords.update(extra_stopwords)
    
    # 先分割句子,保留原始标点用于边界识别
    sentences = sent_tokenize(text)

    filtered_tokens = []
    for sentence in sentences:
        # 对单个句子移除标点、转小写
        sentence = sentence.translate(str.maketrans('', '', string.punctuation))
        sentence = sentence.lower()
        # 分词并过滤停用词
        tokens = word_tokenize(sentence)
        filtered_sentence = [word for word in tokens if word not in bg_stopwords]
        filtered_tokens.append(filtered_sentence)

    return filtered_tokens

内容的提问来源于stack exchange,提问作者mark-de

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 19:57:01