Python生成n-gram和n-1 gram时n-1 gram计数错误调试求助
调试n-1 Gram统计错误的问题
我完全懂你的困扰——统计n-1 gram(也就是历史语境)时,E的计数从预期的3变成了1,明显是代码在遍历或统计环节漏掉了部分元素。我来帮你拆解常见问题点和修复方案:
常见错误原因及修复方案
这种问题大概率出在n-gram的遍历范围或者history序列的截取逻辑上,我结合你的预期输出举个具体例子:
假设你的输入语料是类似这样的(匹配你给出的预期计数):
corpus = ["D", "A", "D", "C", "D", "C", "D", "B", "A", "E", "E", "E"] n = 2 # 以bigram为例,n-1 gram就是单个词
错误的统计方式(导致E计数为1)
如果你的遍历范围写错了,比如写成range(len(corpus)-n)而不是正确的range(len(corpus)-n+1),就会少遍历最后几个n-gram,直接导致E的计数少2次:
from collections import defaultdict def count_ngrams(corpus, n): ngrams = defaultdict(int) history_counts = defaultdict(int) # 错误:遍历范围少了1次,漏掉最后一个bigram for i in range(len(corpus) - n): ngram = tuple(corpus[i:i+n]) ngrams[ngram] += 1 # 取bigram的第一个词作为history history_counts[corpus[i]] += 1 return ngrams, history_counts
调用这段代码后,history_counts里的E只会统计1次,和你遇到的问题完全一致。
修复后的代码
把遍历范围修正为range(len(corpus) - n + 1),确保所有n-gram都被生成并统计:
from collections import defaultdict def count_ngrams(corpus, n): ngrams = defaultdict(int) history_counts = defaultdict(int) # 正确遍历范围:包含所有可能的n-gram for i in range(len(corpus) - n + 1): ngram = tuple(corpus[i:i+n]) ngrams[ngram] += 1 # 以bigram为例,取第一个词作为history history_counts[corpus[i]] += 1 return ngrams, history_counts # 测试 corpus = ["D", "A", "D", "C", "D", "C", "D", "B", "A", "E", "E", "E"] n = 2 _, history_counts = count_ngrams(corpus, n) print(history_counts) # 输出: {'D':4, 'A':2, 'C':3, 'B':1, 'E':3},完全符合预期
另一种可能:history的定义逻辑错误
如果你的history是指n-gram的最后n-1个元素(比如trigram的后两个词作为history),那你需要调整截取逻辑:
# 替换原有的history统计行,取n-gram的后n-1个元素 history = tuple(corpus[i+1:i+n]) history_counts[history[0]] += 1 # 以n=2为例,取后1个元素
总结
你可以优先检查这两个核心问题:
- 遍历n-gram的范围是否正确:必须用
range(len(corpus) - n + 1),避免漏掉最后几个序列 - history的截取逻辑是否匹配需求:确认是取n-gram的前n-1个元素还是后n-1个元素,不要搞混截取位置
如果还是有问题,可以把你的具体代码贴出来,我再帮你进一步排查!
内容的提问来源于stack exchange,提问作者user3320097
相关产品推荐
相关产品推荐

