You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何该语法可在Earley解析器运行却无法适配LALR(1)解析器?

问题:LALR(1)解析器适配补丁语法失败

补丁文件背景与格式说明

这套语法用于定义Verilog RTL代码的简易补丁流程:上游行为代码工程师因IP限制无法知晓库细节,物理设计工程师会将部分行为代码转为结构化门级代码做优化,这类补丁无法合并到上游,最终目标是用这套语法彻底替代现有补丁流程。

工程师提交的补丁文件格式如下:

{ignored space}
//find_start {possible comment or directive}
{content to find}
//find_end
{ignored space}
//replace_start
{content to replace it with}
//replace_end
{ignored_space}
...

当前可用的Earley语法

以下语法在Lark的Earley解析器下可正常生成预期解析树:

start: block+

block: ignore? find_block ignore? replace_block ignore?

find_block:  "//find_start" [comment] content "//find_end" 
replace_block: "//replace_start" content "//replace_end"

_NL: /\n/+
LINE.-10: /.+/
COMMENT: /\s.+/
content: (LINE _NL)+
ignore: (LINE _NL)+
comment: (COMMENT _NL)

%import common.NEWLINE
%import common.WS
%ignore WS
%ignore NEWLINE

LALR(1)解析报错信息

切换到LALR(1)解析器时,抛出如下异常:

lark.exceptions.UnexpectedToken: Unexpected token Token('LINE', 'a') at line 3, column 1.
Expected one of: 
        * _NL
Previous tokens: [Token('LINE', 'asdf')]

测试用示例文本

//find_start asdf
a
//find_end

//replace_start
A
//replace_end

//find_start
b
//find_end
this should be ignored
//replace_start
B
//replace_end

//find_start
c
//find_end

//replace_start
C
C
//replace_end

语法改进建议

原语法核心问题是token歧义和规则边界模糊,LALR(1)作为前瞻1个token的解析器,无法区分LINE应该归属content还是ignore,也无法处理COMMENT和LINE的优先级冲突。以下是适配LALR(1)的优化方案:

1. 明确token优先级与定义

优先匹配指令关键词,避免被通用LINE规则捕获;重新定义各类行的token,消除歧义:

  • 把//find_start、//find_end、//replace_start、//replace_end定义为单独的指令token,优先级高于LINE
  • 重新定义COMMENT_LINE为//find_start后的参数行,而非通用空白开头的行
  • 用负向前瞻定义LINE,排除指令行,明确区分普通行和指令行

2. 重构语法规则,消除上下文歧义

让content和ignore的边界完全由上下文(指令块)决定,避免模糊的规则匹配:

修改后的语法示例:

start: block+

block: ignore? find_block ignore? replace_block ignore?

// 明确指令块结构,绑定注释到find_start后
find_block: FIND_START [COMMENT_LINE] CONTENT_BLOCK FIND_END
replace_block: REPLACE_START CONTENT_BLOCK REPLACE_END

// 定义指令token,优先匹配
FIND_START: "//find_start"
FIND_END: "//find_end"
REPLACE_START: "//replace_start"
REPLACE_END: "//replace_end"

// 定义各类行token
COMMENT_LINE: /.+/  // 匹配//find_start后面的整行注释
LINE: /(?!\/\/(find|replace)_(start|end)).+/  // 匹配非指令行
_NL: /\n/+

// CONTENT_BLOCK是指令块内的内容行
CONTENT_BLOCK: (LINE _NL)+
// ignore是指令块外的任意行
ignore: (LINE _NL)*

%import common.WS
%ignore WS

优化说明

  • 指令token优先匹配,确保//find_start这类行不会被LINE捕获
  • LINE的负向前瞻规则,明确排除指令行,消除token歧义
  • CONTENT_BLOCK和ignore的规则完全由指令块的开始/结束上下文决定,LALR(1)可通过前瞻token明确判断当前行归属
  • 去掉原语法中%ignore NEWLINE设置,改用显式_NL处理换行,避免和规则内的换行逻辑冲突

内容的提问来源于stack exchange,提问作者Ryan Brothers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 02:27:07