You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决使用NLTK生成随机诗歌代码的语法覆盖ValueError

问题:用NLTK的CFG生成随机诗歌时触发ValueError

我是编程写作课程的初学者,正在为期末项目编写生成随机诗歌的Python代码。代码使用NLTK库,通过定义CFG(上下文无关文法)生成句子,但每次运行都会抛出ValueError,提示随机选取的单词不在文法覆盖范围内(如示例中的'unimprovedness')。

代码:

import nltk
nltk.download('words')
nltk.download('punkt')
from nltk.corpus import words
from nltk.tokenize import word_tokenize
from nltk.grammar import CFG
from random import choice

# Choose a random word from the English dictionary
chosen_word = choice(words.words())

# Define the context free grammar
grammar = CFG.fromstring("""
S -> NP VP
NP -> Det N
VP -> V NP
Det -> 'the' | 'a'
N -> 'cat' | 'dog' | 'bird' | 'tree' | 'flower' | 'chosen_word' | 'Fabrikoid'
V -> 'sings' | 'walks' | 'flies' | 'grows' | 'blooms'
""")

# Create a parser object
parser = nltk.ChartParser(grammar)

# Generate a random sentence using the chosen word
sentence = ""
while not sentence:
    # Generate a parse tree for the chosen word
    trees = list(parser.parse(word_tokenize(chosen_word)))

    # Choose a random parse tree
    tree = choice(trees)

    # Generate a sentence from the parse tree
    sentence = tree.label()
    for subtree in tree.subtrees():
        if subtree.label() in ["NP", "VP"]:
            sentence += " " + " ".join(subtree.leaves())

# Print the generated poem
print("Here is your poem:")
print(sentence)

报错信息:

ValueError                                
Traceback (most recent call last)
<ipython-input-57-d6c4184eb90e> in <cell line: 27>()
     27 while not sentence:
     28     # Generate a parse tree for the chosen word
---> 29     trees = list(parser.parse(word_tokenize(chosen_word)))
     30 
     31     # Choose a random parse tree

2 frames
/usr/local/lib/python3.10/dist-packages/nltk/grammar.py in check_coverage(self, tokens)
    663         if missing:
    664             missing = ", ".join(f"{w!r}" for w in missing)
---> 665             raise ValueError(
    666                 "Grammar does not cover some of the " "input words: %r." % missing
    667             )

ValueError: Grammar does not cover some of the input words: "'unimprovedness'".

解决方案

问题根源

  1. 文法定义错误:CFG规则里的'chosen_word'是字符串字面量,不是引用变量chosen_word的实际值,导致文法根本没包含随机选出的单词。
  2. 逻辑颠倒:ChartParser.parse()是用来验证输入句子是否符合文法的工具,不是生成句子的方法,你现在的逻辑是拿随机单词去匹配文法,完全搞反了。
  3. 词性不匹配:从字典选的单词可能不是名词,但文法只把随机单词归为名词N,就算解决字面量问题,遇到其他词性的单词还是会报错。

修正步骤

步骤1:动态将随机名词整合进CFG

先筛选出随机名词(避免词性不兼容),再用字符串拼接生成包含该单词的文法规则:

# 下载词性标注工具
nltk.download('averaged_perceptron_tagger')
from nltk.tag import pos_tag

# 循环选出一个名词(NN/NNS等词性标记为名词)
chosen_word = None
while not chosen_word:
    word = choice(words.words())
    if pos_tag([word])[0][1].startswith('NN'):
        chosen_word = word

# 用f-string动态生成文法
grammar_str = f"""
S -> NP VP
NP -> Det N
VP -> V NP
Det -> 'the' | 'a'
N -> 'cat' | 'dog' | 'bird' | 'tree' | 'flower' | '{chosen_word}' | 'Fabrikoid'
V -> 'sings' | 'walks' | 'flies' | 'grows' | 'blooms'
"""
grammar = CFG.fromstring(grammar_str)

步骤2:用NLTK的生成工具创建句子

使用generate()方法根据文法生成符合规则的句子,而不是用解析工具:

from nltk.parse.generate import generate

# 生成5个候选句子,随机选一个
sentences = list(generate(grammar, n=5))
chosen_sentence = ' '.join(choice(sentences))

修正后的完整代码

import nltk
nltk.download('words')
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')
from nltk.corpus import words
from nltk.tag import pos_tag
from nltk.grammar import CFG
from nltk.parse.generate import generate
from random import choice

# 筛选随机名词
chosen_word = None
while not chosen_word:
    word = choice(words.words())
    if pos_tag([word])[0][1].startswith('NN'):
        chosen_word = word

# 动态构建包含随机名词的CFG
grammar_str = f"""
S -> NP VP
NP -> Det N
VP -> V NP
Det -> 'the' | 'a'
N -> 'cat' | 'dog' | 'bird' | 'tree' | 'flower' | '{chosen_word}' | 'Fabrikoid'
V -> 'sings' | 'walks' | 'flies' | 'grows' | 'blooms'
"""
grammar = CFG.fromstring(grammar_str)

# 生成随机诗歌句子
sentences = list(generate(grammar, n=5))
poem_sentence = ' '.join(choice(sentences))

# 输出结果
print("Here is your poem:")
print(poem_sentence)

额外说明

  • 词性筛选可以确保随机单词能匹配文法中的N规则,避免无效单词导致的报错。
  • 调整generate()的n参数可以生成更多候选句子,增加诗歌的随机性。
  • 如果想要更复杂的诗歌结构,可以扩展CFG规则,比如加入形容词、副词,或者多句子组合的规则。

内容的提问来源于stack exchange,提问作者IceBulb v2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 21:30:32