为何该语法可在Earley解析器运行却无法适配LALR(1)解析器?
问题:LALR(1)解析器适配补丁语法失败
补丁文件背景与格式说明
这套语法用于定义Verilog RTL代码的简易补丁流程:上游行为代码工程师因IP限制无法知晓库细节,物理设计工程师会将部分行为代码转为结构化门级代码做优化,这类补丁无法合并到上游,最终目标是用这套语法彻底替代现有补丁流程。
工程师提交的补丁文件格式如下:
{ignored space} //find_start {possible comment or directive} {content to find} //find_end {ignored space} //replace_start {content to replace it with} //replace_end {ignored_space} ...
当前可用的Earley语法
以下语法在Lark的Earley解析器下可正常生成预期解析树:
start: block+ block: ignore? find_block ignore? replace_block ignore? find_block: "//find_start" [comment] content "//find_end" replace_block: "//replace_start" content "//replace_end" _NL: /\n/+ LINE.-10: /.+/ COMMENT: /\s.+/ content: (LINE _NL)+ ignore: (LINE _NL)+ comment: (COMMENT _NL) %import common.NEWLINE %import common.WS %ignore WS %ignore NEWLINE
LALR(1)解析报错信息
切换到LALR(1)解析器时,抛出如下异常:
lark.exceptions.UnexpectedToken: Unexpected token Token('LINE', 'a') at line 3, column 1. Expected one of: * _NL Previous tokens: [Token('LINE', 'asdf')]
测试用示例文本
//find_start asdf a //find_end //replace_start A //replace_end //find_start b //find_end this should be ignored //replace_start B //replace_end //find_start c //find_end //replace_start C C //replace_end
语法改进建议
原语法核心问题是token歧义和规则边界模糊,LALR(1)作为前瞻1个token的解析器,无法区分LINE应该归属content还是ignore,也无法处理COMMENT和LINE的优先级冲突。以下是适配LALR(1)的优化方案:
1. 明确token优先级与定义
优先匹配指令关键词,避免被通用LINE规则捕获;重新定义各类行的token,消除歧义:
- 把
//find_start、//find_end、//replace_start、//replace_end定义为单独的指令token,优先级高于LINE - 重新定义
COMMENT_LINE为//find_start后的参数行,而非通用空白开头的行 - 用负向前瞻定义
LINE,排除指令行,明确区分普通行和指令行
2. 重构语法规则,消除上下文歧义
让content和ignore的边界完全由上下文(指令块)决定,避免模糊的规则匹配:
修改后的语法示例:
start: block+ block: ignore? find_block ignore? replace_block ignore? // 明确指令块结构,绑定注释到find_start后 find_block: FIND_START [COMMENT_LINE] CONTENT_BLOCK FIND_END replace_block: REPLACE_START CONTENT_BLOCK REPLACE_END // 定义指令token,优先匹配 FIND_START: "//find_start" FIND_END: "//find_end" REPLACE_START: "//replace_start" REPLACE_END: "//replace_end" // 定义各类行token COMMENT_LINE: /.+/ // 匹配//find_start后面的整行注释 LINE: /(?!\/\/(find|replace)_(start|end)).+/ // 匹配非指令行 _NL: /\n/+ // CONTENT_BLOCK是指令块内的内容行 CONTENT_BLOCK: (LINE _NL)+ // ignore是指令块外的任意行 ignore: (LINE _NL)* %import common.WS %ignore WS
优化说明
- 指令token优先匹配,确保
//find_start这类行不会被LINE捕获 LINE的负向前瞻规则,明确排除指令行,消除token歧义CONTENT_BLOCK和ignore的规则完全由指令块的开始/结束上下文决定,LALR(1)可通过前瞻token明确判断当前行归属- 去掉原语法中
%ignore NEWLINE设置,改用显式_NL处理换行,避免和规则内的换行逻辑冲突
内容的提问来源于stack exchange,提问作者Ryan Brothers
相关产品推荐
相关产品推荐

