Lark语法优先级问题求助:解析结果与预期不符
The issue you're facing boils down to how PEG parsers (like Lark) prioritize rule matching—they try rules in the order they’re listed and greedily lock into the first valid match. In your original code, the set rule can already match the entire input "3xfoo" (since it includes an optional nb prefix), so the parser never gets the chance to split the nb and set into separate nodes via the set_mul rule.
Here's a straightforward fix:
Adjust Rule Order and Structure
Reorder the start rule to try set_mul first, and define set_mul as explicitly requiring an nb followed by a set. This ensures the parser prioritizes splitting the input into nb and set when possible, before falling back to the standalone set rule.
from lark import Lark parser = Lark(r""" start: set_mul | set set_mul: nb set set: [nb] "foo" # Keep this if you still need `set` to accept an optional nb prefix nb: INT "x" %import common.INT %import common.WS %ignore WS """, start='start') input_str = "3xfoo" p = parser.parse(input_str) print(p.pretty())
Expected Output
start set_mul nb 3 set
Why This Works
- Rule Order: By placing
set_mulbeforesetin thestartrule, the parser first attempts to matchnb + set. For"3xfoo", it successfully matchesnb(3x) and thenset(foo), so it uses this interpretation instead of falling back to thesetrule. - Explicit
set_mulDefinition: Definingset_mulasnb set(rather than[nb] set) removes ambiguity—now it only matches when there’s annbpreceding aset, which aligns perfectly with your desired AST structure.
Alternative (If You Don’t Need set to Accept nb)
If you don’t actually require set to have an optional nb prefix, you can simplify the rules further to eliminate all ambiguity entirely:
from lark import Lark parser = Lark(r""" start: set_mul set_mul: nb set | set set: "foo" nb: INT "x" %import common.INT %import common.WS %ignore WS """, start='start')
This version will still produce your desired output for "3xfoo" and cleanly handle standalone "foo" inputs as well.
内容的提问来源于stack exchange,提问作者Kypaz

