使用NLTK trigram模型生成句子时遇报错问题求助
解决NLTK Trigram语言建模中的两个关键错误
第一个错误:AttributeError: 'int' object has no attribute 'items'
原因
d是defaultdict(Counter)类型,意味着d[a,b]是一个Counter对象。当你执行d[a,b] += freq_tri[a,b,c]时,Counter的+=操作要求右侧是另一个Counter(或键值对迭代器),但freq_tri[a,b,c]返回的是该trigram的出现次数(整数),类型不匹配导致报错。
修正写法
把整数次数加到Counter对应键的计数上,二选一即可:
# 直接通过键赋值累加 d[a, b][c] += freq_tri[(a, b, c)]
或者:
# 使用Counter的update方法 d[a, b].update({c: freq_tri[(a, b, c)]})
第二个错误:ValueError: not enough values to unpack (expected 3, got 2)
原因
FreqDist.items()返回的是键值对元组,每个元组结构是(trigram_tuple, count):第一个元素是trigram的三元组(比如(a,b,c)),第二个是该trigram的出现次数。你用for a,b,c in freq_tri.items()时,每个迭代元素只有两个值,无法拆成三个变量,因此报错。
修正循环写法
先拆分键值对,再从三元组里提取a、b、c:
for trigram_tuple, count in freq_tri.items(): a, b, c = trigram_tuple if a is not None and b is not None: d[a, b][c] += count
直接用count变量比再次查询freq_tri更高效。
代码中其他必须修正的问题
- 循环缩进错误
原代码中for sentence in sents:后没有缩进,导致只处理了最后一个句子,所有n-gram都来自单一句子,完全不符合语言建模需求。修正后缩进处理每个句子:
for sentence in sents: sentence = list(map(lambda x: x.lower(), sentence)) cleaned_sentence = [] for word in sentence: if word != '.': cleaned_sentence.append(word) unigram.append(word) tokenized_text.append(cleaned_sentence) bigram.extend(list(ngrams(cleaned_sentence, 2, pad_left=True, pad_right=True))) trigram.extend(list(ngrams(cleaned_sentence, 3, pad_left=True, pad_right=True))) fourgram.extend(list(ngrams(cleaned_sentence, 4, pad_left=True, pad_right=True)))
注:原代码直接修改原sentence(sentence.remove(word))会导致循环跳过元素,改用新列表收集更安全。
remove_stopwords函数逻辑错误
原函数的判断逻辑是“只要n-gram中有一个词不在移除列表就保留”,这会保留包含停用词/标点的n-gram,不符合需求。正确逻辑是n-gram中所有词都不在移除列表才保留:
def remove_stopwords(ngrams_list): filtered = [] for gram in ngrams_list: valid = True for word in gram: # 跳过pad生成的None值 if word is not None and word in removal_list: valid = False break if valid: filtered.append(gram) return filtered
另外,原代码中unigram = remove_stopwords(unigram)是错误的(unigram是单个词的列表,函数处理的是元组类型的n-gram),单独过滤unigram:
unigram = [word for word in unigram if word not in removal_list]
修正后的核心代码片段
# ...(import和NLTK资源下载部分不变) # 生成unigrams、bigrams、trigrams unigram = [] bigram = [] trigram = [] fourgram = [] tokenized_text = [] for sentence in sents: sentence = list(map(lambda x: x.lower(), sentence)) cleaned_sentence = [] for word in sentence: if word != '.': cleaned_sentence.append(word) unigram.append(word) tokenized_text.append(cleaned_sentence) bigram.extend(list(ngrams(cleaned_sentence, 2, pad_left=True, pad_right=True))) trigram.extend(list(ngrams(cleaned_sentence, 3, pad_left=True, pad_right=True))) fourgram.extend(list(ngrams(cleaned_sentence, 4, pad_left=True, pad_right=True))) # 过滤unigram unigram = [word for word in unigram if word not in removal_list] # 过滤n-grams def remove_stopwords(ngrams_list): filtered = [] for gram in ngrams_list: valid = True for word in gram: if word is not None and word in removal_list: valid = False break if valid: filtered.append(gram) return filtered bigram = remove_stopwords(bigram) trigram = remove_stopwords(trigram) fourgram = remove_stopwords(fourgram) # 生成频率分布 freq_bi = FreqDist(bigram) freq_tri = FreqDist(trigram) freq_four = FreqDist(fourgram) d = defaultdict(Counter) # 正确构建trigram条件概率映射 for trigram_tuple, count in freq_tri.items(): a, b, c = trigram_tuple if a is not None and b is not None: d[a, b][c] += count # ...(pick_word函数和句子生成部分不变)
内容的提问来源于stack exchange,提问作者Jesper Ezra
相关产品推荐
相关产品推荐

