如何随机交换CFG文法规则?含报错排查与实现需求
CFG规则随机交换实现及报错修复
报错修复
你遇到的SyntaxError是因为substitute库基于Python2语法编写,其中lambda (a,b):a==b的元组参数解包在Python3中已被移除。直接删除import substitute语句即可,我们将手动实现规则交换逻辑。
非终结符规则随机交换思路
对任意非终结符(如S、NP、VP),按以下步骤随机交换其规则分支:
- 获取该非终结符对应的所有产生式
- 提取产生式右侧的所有候选分支
- 随机打乱分支顺序,生成新的产生式
- 用新产生式替换原文法中的对应规则
修正后的完整代码
import spacy import nltk from nltk.parse.generate import generate import random from nltk.grammar import CFG from nltk import Nonterminal nlp = spacy.load("en_core_web_sm") fin = [ "Wow!The movie was a complete joy to watch, with an incredible cast delivering fantastic performances. The special effects were stunning. I highly recommend this movie to everyone." ] def shuffle_nt_productions(grammar, nt): """随机打乱指定非终结符的产生式分支顺序""" # 筛选目标非终结符的所有产生式 orig_productions = [prod for prod in grammar.productions() if prod.lhs() == nt] if not orig_productions: return grammar # 提取并打乱所有分支 rhs_list = [prod.rhs() for prod in orig_productions] random.shuffle(rhs_list) # 构建新产生式 new_productions = [CFG.production(nt, rhs) for rhs in rhs_list] # 生成新文法:保留其他规则,替换目标非终结符的规则 other_productions = [prod for prod in grammar.productions() if prod.lhs() != nt] return CFG(grammar.start(), other_productions + new_productions) for line in fin: sent = line words = [x.lower() for x in nltk.word_tokenize(sent)] sent = ' '.join(words) doc = nlp(sent) noun = [] verb = [] adj = [] det = [] adv = [] not_or_no = "" # 词性标注提取 for token in doc: if token.text in ("not", "no"): not_or_no = token.text continue token_text = f"{not_or_no} {token.text}" if not_or_no else token.text if token.pos_ in ("NOUN", "PROPN", "PRON"): noun.append(token_text) elif token.pos_ in ("VERB", "AUX"): verb.append(token_text) elif token.pos_ == "DET": det.append(token.text) elif token.pos_ == "ADJ": adj.append(token_text) elif token.pos_ == "ADV": adv.append(token_text) not_or_no = "" # 构建终端符号选项字符串 def build_symbol_str(items): return " | ".join([f'"{item}"' for item in items]) if items else "" NOUN = build_symbol_str(noun) VERB = build_symbol_str(verb) DET = build_symbol_str(det) ADJ = build_symbol_str(adj) ADV = build_symbol_str(adv) # 初始化CFG文法 grammar = CFG.fromstring(f""" S -> NP VP | VP NP NP -> DET N | ADJ N VP -> V NP | V ADJ | V ADV V -> {VERB} N -> {NOUN} ADJ -> {ADJ} DET -> {DET} ADV -> {ADV} """) # 生成50个随机交换规则的句子 for _ in range(50): # 随机选择要交换规则的非终结符(可单选或多选) target_nts = random.choice([ [Nonterminal('S')], [Nonterminal('NP')], [Nonterminal('VP')], [Nonterminal('S'), Nonterminal('NP')], [Nonterminal('S'), Nonterminal('VP')] ]) new_grammar = grammar for nt in target_nts: new_grammar = shuffle_nt_productions(new_grammar, nt) # 生成并打印句子 for sentence in generate(new_grammar, n=1): print(' '.join(sentence))
关键说明
shuffle_nt_productions:核心函数,负责对指定非终结符的规则分支进行随机打乱,生成新文法。- 词性提取逻辑优化:简化了终端符号字符串的构建,避免冗余代码。
- 规则交换多样性:支持随机选择单个或多个非终结符进行规则交换,生成更多样化的句子。
内容的提问来源于stack exchange,提问作者Sunjaree
相关产品推荐
相关产品推荐

