如何解决使用NLTK生成随机诗歌代码的语法覆盖ValueError
问题:用NLTK的CFG生成随机诗歌时触发ValueError
我是编程写作课程的初学者,正在为期末项目编写生成随机诗歌的Python代码。代码使用NLTK库,通过定义CFG(上下文无关文法)生成句子,但每次运行都会抛出ValueError,提示随机选取的单词不在文法覆盖范围内(如示例中的'unimprovedness')。
代码:
import nltk nltk.download('words') nltk.download('punkt') from nltk.corpus import words from nltk.tokenize import word_tokenize from nltk.grammar import CFG from random import choice # Choose a random word from the English dictionary chosen_word = choice(words.words()) # Define the context free grammar grammar = CFG.fromstring(""" S -> NP VP NP -> Det N VP -> V NP Det -> 'the' | 'a' N -> 'cat' | 'dog' | 'bird' | 'tree' | 'flower' | 'chosen_word' | 'Fabrikoid' V -> 'sings' | 'walks' | 'flies' | 'grows' | 'blooms' """) # Create a parser object parser = nltk.ChartParser(grammar) # Generate a random sentence using the chosen word sentence = "" while not sentence: # Generate a parse tree for the chosen word trees = list(parser.parse(word_tokenize(chosen_word))) # Choose a random parse tree tree = choice(trees) # Generate a sentence from the parse tree sentence = tree.label() for subtree in tree.subtrees(): if subtree.label() in ["NP", "VP"]: sentence += " " + " ".join(subtree.leaves()) # Print the generated poem print("Here is your poem:") print(sentence)
报错信息:
ValueError Traceback (most recent call last) <ipython-input-57-d6c4184eb90e> in <cell line: 27>() 27 while not sentence: 28 # Generate a parse tree for the chosen word ---> 29 trees = list(parser.parse(word_tokenize(chosen_word))) 30 31 # Choose a random parse tree 2 frames /usr/local/lib/python3.10/dist-packages/nltk/grammar.py in check_coverage(self, tokens) 663 if missing: 664 missing = ", ".join(f"{w!r}" for w in missing) ---> 665 raise ValueError( 666 "Grammar does not cover some of the " "input words: %r." % missing 667 ) ValueError: Grammar does not cover some of the input words: "'unimprovedness'".
解决方案
问题根源
- 文法定义错误:CFG规则里的
'chosen_word'是字符串字面量,不是引用变量chosen_word的实际值,导致文法根本没包含随机选出的单词。 - 逻辑颠倒:
ChartParser.parse()是用来验证输入句子是否符合文法的工具,不是生成句子的方法,你现在的逻辑是拿随机单词去匹配文法,完全搞反了。 - 词性不匹配:从字典选的单词可能不是名词,但文法只把随机单词归为名词
N,就算解决字面量问题,遇到其他词性的单词还是会报错。
修正步骤
步骤1:动态将随机名词整合进CFG
先筛选出随机名词(避免词性不兼容),再用字符串拼接生成包含该单词的文法规则:
# 下载词性标注工具 nltk.download('averaged_perceptron_tagger') from nltk.tag import pos_tag # 循环选出一个名词(NN/NNS等词性标记为名词) chosen_word = None while not chosen_word: word = choice(words.words()) if pos_tag([word])[0][1].startswith('NN'): chosen_word = word # 用f-string动态生成文法 grammar_str = f""" S -> NP VP NP -> Det N VP -> V NP Det -> 'the' | 'a' N -> 'cat' | 'dog' | 'bird' | 'tree' | 'flower' | '{chosen_word}' | 'Fabrikoid' V -> 'sings' | 'walks' | 'flies' | 'grows' | 'blooms' """ grammar = CFG.fromstring(grammar_str)
步骤2:用NLTK的生成工具创建句子
使用generate()方法根据文法生成符合规则的句子,而不是用解析工具:
from nltk.parse.generate import generate # 生成5个候选句子,随机选一个 sentences = list(generate(grammar, n=5)) chosen_sentence = ' '.join(choice(sentences))
修正后的完整代码
import nltk nltk.download('words') nltk.download('punkt') nltk.download('averaged_perceptron_tagger') from nltk.corpus import words from nltk.tag import pos_tag from nltk.grammar import CFG from nltk.parse.generate import generate from random import choice # 筛选随机名词 chosen_word = None while not chosen_word: word = choice(words.words()) if pos_tag([word])[0][1].startswith('NN'): chosen_word = word # 动态构建包含随机名词的CFG grammar_str = f""" S -> NP VP NP -> Det N VP -> V NP Det -> 'the' | 'a' N -> 'cat' | 'dog' | 'bird' | 'tree' | 'flower' | '{chosen_word}' | 'Fabrikoid' V -> 'sings' | 'walks' | 'flies' | 'grows' | 'blooms' """ grammar = CFG.fromstring(grammar_str) # 生成随机诗歌句子 sentences = list(generate(grammar, n=5)) poem_sentence = ' '.join(choice(sentences)) # 输出结果 print("Here is your poem:") print(poem_sentence)
额外说明
- 词性筛选可以确保随机单词能匹配文法中的
N规则,避免无效单词导致的报错。 - 调整
generate()的n参数可以生成更多候选句子,增加诗歌的随机性。 - 如果想要更复杂的诗歌结构,可以扩展CFG规则,比如加入形容词、副词,或者多句子组合的规则。
内容的提问来源于stack exchange,提问作者IceBulb v2
相关产品推荐
相关产品推荐

