You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何随机交换CFG文法规则?含报错排查与实现需求

CFG规则随机交换实现及报错修复

报错修复

你遇到的SyntaxError是因为substitute库基于Python2语法编写,其中lambda (a,b):a==b的元组参数解包在Python3中已被移除。直接删除import substitute语句即可,我们将手动实现规则交换逻辑。

非终结符规则随机交换思路

对任意非终结符(如S、NP、VP),按以下步骤随机交换其规则分支:

  • 获取该非终结符对应的所有产生式
  • 提取产生式右侧的所有候选分支
  • 随机打乱分支顺序,生成新的产生式
  • 用新产生式替换原文法中的对应规则

修正后的完整代码

import spacy
import nltk
from nltk.parse.generate import generate
import random
from nltk.grammar import CFG
from nltk import Nonterminal

nlp = spacy.load("en_core_web_sm")

fin = [
    "Wow!The movie was a complete joy to watch, with an incredible cast delivering fantastic performances. The special effects were stunning. I highly recommend this movie to everyone."
]  

def shuffle_nt_productions(grammar, nt):
    """随机打乱指定非终结符的产生式分支顺序"""
    # 筛选目标非终结符的所有产生式
    orig_productions = [prod for prod in grammar.productions() if prod.lhs() == nt]
    if not orig_productions:
        return grammar
    
    # 提取并打乱所有分支
    rhs_list = [prod.rhs() for prod in orig_productions]
    random.shuffle(rhs_list)
    
    # 构建新产生式
    new_productions = [CFG.production(nt, rhs) for rhs in rhs_list]
    
    # 生成新文法:保留其他规则,替换目标非终结符的规则
    other_productions = [prod for prod in grammar.productions() if prod.lhs() != nt]
    return CFG(grammar.start(), other_productions + new_productions)

for line in fin:
    sent = line
    words = [x.lower() for x in nltk.word_tokenize(sent)]
    sent = ' '.join(words)
    doc = nlp(sent)

    noun = []
    verb = []
    adj = []
    det = []
    adv = []
    not_or_no = ""

    # 词性标注提取
    for token in doc:
        if token.text in ("not", "no"):
            not_or_no = token.text
            continue

        token_text = f"{not_or_no} {token.text}" if not_or_no else token.text
        if token.pos_ in ("NOUN", "PROPN", "PRON"):
            noun.append(token_text)
        elif token.pos_ in ("VERB", "AUX"):
            verb.append(token_text)
        elif token.pos_ == "DET":
            det.append(token.text)
        elif token.pos_ == "ADJ":
            adj.append(token_text)
        elif token.pos_ == "ADV":
            adv.append(token_text)
        
        not_or_no = ""

    # 构建终端符号选项字符串
    def build_symbol_str(items):
        return " | ".join([f'"{item}"' for item in items]) if items else ""

    NOUN = build_symbol_str(noun)
    VERB = build_symbol_str(verb)
    DET = build_symbol_str(det)
    ADJ = build_symbol_str(adj)
    ADV = build_symbol_str(adv)

    # 初始化CFG文法
    grammar = CFG.fromstring(f"""
    S -> NP VP | VP NP
    NP -> DET N | ADJ N 
    VP -> V NP | V ADJ | V ADV
    V -> {VERB}
    N -> {NOUN}
    ADJ -> {ADJ}
    DET -> {DET}
    ADV -> {ADV}
    """)

    # 生成50个随机交换规则的句子
    for _ in range(50):
        # 随机选择要交换规则的非终结符(可单选或多选)
        target_nts = random.choice([
            [Nonterminal('S')],
            [Nonterminal('NP')],
            [Nonterminal('VP')],
            [Nonterminal('S'), Nonterminal('NP')],
            [Nonterminal('S'), Nonterminal('VP')]
        ])
        new_grammar = grammar
        for nt in target_nts:
            new_grammar = shuffle_nt_productions(new_grammar, nt)
        
        # 生成并打印句子
        for sentence in generate(new_grammar, n=1):
            print(' '.join(sentence))

关键说明

  • shuffle_nt_productions:核心函数,负责对指定非终结符的规则分支进行随机打乱,生成新文法。
  • 词性提取逻辑优化:简化了终端符号字符串的构建,避免冗余代码。
  • 规则交换多样性:支持随机选择单个或多个非终结符进行规则交换,生成更多样化的句子。

内容的提问来源于stack exchange,提问作者Sunjaree

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 21:05:21