You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何减少Lark语法中NEWLINE/DOUBLE_NEWLINE相关规则的重复定义?

减少Lark语法中换行相关规则重复的方案

你的核心需求是用双换行(\n\n)作为叶倾斜树的分隔标记(替代示例中的|),同时避免为单/双换行重复定义几乎一致的规则,还要保证Indenter不对双换行执行退缩进操作。以下是可行的优化方案:

优化思路

通过统一规则逻辑和调整Indenter的令牌处理逻辑,彻底消除规则重复:

  • 用单一规则组处理“节点/子树 + 可选分隔符”的结构
  • 让Indenter仅处理单换行,忽略双换行的缩进干扰

完整实现代码

from lark import Lark
from lark.indenter import Indenter

grammar = """\
start: tree
tree: tree_part+
tree_part: (node+ | sub_tree) [DNL]
?node: NAME NL

?sub_tree: _INDENT tree _DEDENT

NL: /\\n */
DNL: /\\n\\n */

%declare _INDENT _DEDENT
%import common.CNAME -> NAME
"""

# 对应 [[[a, b], c, d], e, [g], f] 的示例输入
sample_string = """\
a
b

c
d

e
    g
f
"""

class TreeIndenter(Indenter):
    NL_type = 'NL'
    OPEN_PAREN_types = []
    CLOSE_PAREN_types = []
    INDENT_type = '_INDENT'
    DEDENT_type = '_DEDENT'
    tab_len = 4

    def process(self, stream):
        # 仅对单换行NL执行缩进处理,跳过双换行DNL
        for token in stream:
            if token.type == 'DNL':
                yield token
            elif token.type == self.NL_type:
                yield from super().process([token])
            else:
                yield token

parser = Lark(grammar, postlex=TreeIndenter(), debug=True)
for i, t in enumerate(parser.lex(sample_string)):
    print((t.line, t.column), repr(t))
parse_tree = parser.parse(sample_string)
print(parse_tree.pretty())

方案细节说明

  1. 规则简化:

    • 删除了所有重复的tree_d、node_d、sub_tree_d规则,用tree_part统一表示“一组节点/子树 + 可选双换行分隔符”
    • tree由多个tree_part组成,完全匹配你需要的[a b | c d | e [g] f]结构(|对应DNL)
  2. Indenter适配:

    • 重写Indenter的process方法,跳过对DNL令牌的缩进处理,只处理单换行NL
    • 这样既保留了双换行的语法分隔意义,又不会让Indenter错误地对双换行执行退缩进操作
  3. 语法正确性:

    • 示例输入解析后的语法树会自然分组:[a,b]、[c,d]、[e,[g],f],完全符合预期结构

内容的提问来源于stack exchange,提问作者Tom Huntington

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 05:00:56