You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用NLTK trigram模型生成句子时遇报错问题求助

解决NLTK Trigram语言建模中的两个关键错误

第一个错误:AttributeError: 'int' object has no attribute 'items'

原因

d是defaultdict(Counter)类型,意味着d[a,b]是一个Counter对象。当你执行d[a,b] += freq_tri[a,b,c]时,Counter的+=操作要求右侧是另一个Counter(或键值对迭代器),但freq_tri[a,b,c]返回的是该trigram的出现次数(整数),类型不匹配导致报错。

修正写法

把整数次数加到Counter对应键的计数上,二选一即可:

# 直接通过键赋值累加
d[a, b][c] += freq_tri[(a, b, c)]

或者:

# 使用Counter的update方法
d[a, b].update({c: freq_tri[(a, b, c)]})

第二个错误:ValueError: not enough values to unpack (expected 3, got 2)

原因

FreqDist.items()返回的是键值对元组,每个元组结构是(trigram_tuple, count):第一个元素是trigram的三元组(比如(a,b,c)),第二个是该trigram的出现次数。你用for a,b,c in freq_tri.items()时,每个迭代元素只有两个值,无法拆成三个变量,因此报错。

修正循环写法

先拆分键值对,再从三元组里提取a、b、c:

for trigram_tuple, count in freq_tri.items():
    a, b, c = trigram_tuple
    if a is not None and b is not None:
        d[a, b][c] += count

直接用count变量比再次查询freq_tri更高效。

代码中其他必须修正的问题

  1. 循环缩进错误
    原代码中for sentence in sents:后没有缩进,导致只处理了最后一个句子,所有n-gram都来自单一句子,完全不符合语言建模需求。修正后缩进处理每个句子:
for sentence in sents:
    sentence = list(map(lambda x: x.lower(), sentence))
    cleaned_sentence = []
    for word in sentence:
        if word != '.':
            cleaned_sentence.append(word)
            unigram.append(word)
    tokenized_text.append(cleaned_sentence)
    bigram.extend(list(ngrams(cleaned_sentence, 2, pad_left=True, pad_right=True)))
    trigram.extend(list(ngrams(cleaned_sentence, 3, pad_left=True, pad_right=True)))
    fourgram.extend(list(ngrams(cleaned_sentence, 4, pad_left=True, pad_right=True)))

注:原代码直接修改原sentence(sentence.remove(word))会导致循环跳过元素,改用新列表收集更安全。

  1. remove_stopwords函数逻辑错误
    原函数的判断逻辑是“只要n-gram中有一个词不在移除列表就保留”,这会保留包含停用词/标点的n-gram,不符合需求。正确逻辑是n-gram中所有词都不在移除列表才保留:
def remove_stopwords(ngrams_list):
    filtered = []
    for gram in ngrams_list:
        valid = True
        for word in gram:
            # 跳过pad生成的None值
            if word is not None and word in removal_list:
                valid = False
                break
        if valid:
            filtered.append(gram)
    return filtered

另外,原代码中unigram = remove_stopwords(unigram)是错误的(unigram是单个词的列表,函数处理的是元组类型的n-gram),单独过滤unigram:

unigram = [word for word in unigram if word not in removal_list]

修正后的核心代码片段

# ...(import和NLTK资源下载部分不变)

# 生成unigrams、bigrams、trigrams
unigram = []
bigram = []
trigram = []
fourgram = []
tokenized_text = []

for sentence in sents:
    sentence = list(map(lambda x: x.lower(), sentence))
    cleaned_sentence = []
    for word in sentence:
        if word != '.':
            cleaned_sentence.append(word)
            unigram.append(word)
    tokenized_text.append(cleaned_sentence)
    bigram.extend(list(ngrams(cleaned_sentence, 2, pad_left=True, pad_right=True)))
    trigram.extend(list(ngrams(cleaned_sentence, 3, pad_left=True, pad_right=True)))
    fourgram.extend(list(ngrams(cleaned_sentence, 4, pad_left=True, pad_right=True)))

# 过滤unigram
unigram = [word for word in unigram if word not in removal_list]

# 过滤n-grams
def remove_stopwords(ngrams_list):
    filtered = []
    for gram in ngrams_list:
        valid = True
        for word in gram:
            if word is not None and word in removal_list:
                valid = False
                break
        if valid:
            filtered.append(gram)
    return filtered

bigram = remove_stopwords(bigram)
trigram = remove_stopwords(trigram)
fourgram = remove_stopwords(fourgram)

# 生成频率分布
freq_bi = FreqDist(bigram)
freq_tri = FreqDist(trigram)
freq_four = FreqDist(fourgram)

d = defaultdict(Counter)

# 正确构建trigram条件概率映射
for trigram_tuple, count in freq_tri.items():
    a, b, c = trigram_tuple
    if a is not None and b is not None:
        d[a, b][c] += count

# ...(pick_word函数和句子生成部分不变)

内容的提问来源于stack exchange,提问作者Jesper Ezra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 12:20:28