保加利亚语文本生成脚本陷入无限循环问题求助
保加利亚语文本生成器无限循环问题排查
问题描述
开发保加利亚语简易文本生成器时,代码陷入无限循环。调试发现生成文本的嵌套for循环中print语句从未执行,无法向text列表添加新单词。原代码使用reuters.sents()可正常运行,现基于自定义分词和训练文本出现异常。
相关代码
主代码
from tokenization import tokenize_bulgarian_text from nltk import bigrams, trigrams from collections import Counter, defaultdict import random with open('IvanVazov1.txt', 'r', encoding='utf-8') as f: data = f.read() # Tokenize text tokenized_sentences = tokenize_bulgarian_text(data) print(tokenized_sentences) # Create a placeholder for model model = defaultdict(lambda: defaultdict(lambda: 0)) # Count frequency of co-occurance for sentence in tokenized_sentences: for w1, w2, w3 in trigrams(sentence, pad_right=True, pad_left=True): print("Trigram:", (w1, w2, w3)) model[(w1, w2)][w3] += 1 # Transform the counts to probabilities for w1_w2 in model: total_count = float(sum(model[w1_w2].values())) for w3 in model[w1_w2]: model[w1_w2][w3] /= total_count print(model) # starting words text = ["беше"] sentence_finished = False print("Starting words:", text) while not sentence_finished: # select a random probability threshold r = random.random() accumulator = .0 for word in model[tuple(text[-2:])].keys(): # print statement never executes print(f"word: {word}, probability: {model[tuple(text[-2:])][word]}, accumulator: {accumulator}") accumulator += model[tuple(text[-2:])][word] # select words that are above the probability threshold if accumulator >= r: text.append(word) break if text[-2:] == [None, None]: print("End of sentence.") sentence_finished = True print (' '.join([t for t in text if t]))
tokenization.py代码
import nltk from nltk.tokenize import sent_tokenize, word_tokenize import string from spacy.lang.bg.stop_words import STOP_WORDS as bg_stopwords def tokenize_bulgarian_text(text): extra_stopwords = {"—", "“", "„", "не", "та", "па"} bg_stopwords.update(extra_stopwords) # Remove punctuation and lowercase all letters text = text.translate(str.maketrans('', '', string.punctuation)) text = text.lower() # Split the text into individual sentences sentences = sent_tokenize(text) # Tokenize and filter each sentence filtered_tokens = [] for sentence in sentences: # Split the sentence into individual words or tokens tokens = word_tokenize(sentence) # Remove any stop words from the tokenized text filtered_sentence = [word for word in tokens if word not in bg_stopwords] filtered_tokens.append(filtered_sentence) return filtered_tokens
训练文本
Нощта беше влажна и мрачна и браилските улици пустееха. Студената декемврийска мъгла, която обикновено пада покрай бреговете на Дунава, се беше напластила в една от главните улици на града и задушаваше с отровния си дъх последните минувачи, които бързаха да се приберат у дома си.
问题分析与解决方案
1. 核心问题:初始查询键不存在
你的text初始值只有单个单词["беше"],执行model[tuple(text[-2:])]时,得到的是单元素元组("беше",),但你的trigram模型中所有键都是二元组(比如(None, None)、(None, "нощта")、("нощта", "беше")等),因此这个查询会返回一个空的defaultdict,导致嵌套for循环根本不会执行,text永远无法新增单词,最终陷入无限循环。
解决方法:
将初始text设置为训练数据中存在的连续二元组,比如从训练文本中选取["нощта", "беше"],这样text[-2:]就能匹配模型中存在的键("нощта", "беше"),从而正常获取后续单词:
# starting words text = ["нощта", "беше"]
2. 次要问题:分词步骤顺序错误
当前分词函数先全局移除所有标点,再执行sent_tokenize分割句子。但sent_tokenize依赖标点(如句号)识别句子边界,移除标点后会导致句子分割错误,进而影响trigram模型的构建。
解决方法:调整分词步骤顺序,先分割句子,再对每个句子单独处理标点和小写:
def tokenize_bulgarian_text(text): extra_stopwords = {"—", "“", "„", "не", "та", "па"} bg_stopwords.update(extra_stopwords) # 先分割句子,保留原始标点用于边界识别 sentences = sent_tokenize(text) filtered_tokens = [] for sentence in sentences: # 对单个句子移除标点、转小写 sentence = sentence.translate(str.maketrans('', '', string.punctuation)) sentence = sentence.lower() # 分词并过滤停用词 tokens = word_tokenize(sentence) filtered_sentence = [word for word in tokens if word not in bg_stopwords] filtered_tokens.append(filtered_sentence) return filtered_tokens
内容的提问来源于stack exchange,提问作者mark-de
相关产品推荐
相关产品推荐

