如何优化无超前查看(Lookahead)的Grammar语法以解析带单/双竖线的文本?
Let's break down why your simplified grammars aren't working, then fix them with clean, efficient solutions.
Why Your Attempts Failed
The "too-simple" grammar:
grammar HC { token TOP { <pre> <divider> <post> } token pre { \N*? } token divider { <[|]> ** 1..2 } token post { \N* } }The
\N*?non-greedy match will always match 0 characters first (since that's the "least" it can match), leaving<divider>to start parsing from the very beginning of your input. That's why it fails to capture the text before the separator.The repeated single-character token:
grammar HC { token TOP { <pre>* <divider> <post> } token pre { \N } token divider { <[|]> ** 1..2 } token post { \N* } }Using
<pre>*repeats the single-character\Ntoken, so every character before the separator gets its own<pre>node. That's not useful when you want the entire pre-separator text as one capture.
Solution 1: Simple Exclusion (Best for Pre-Text Without |)
If your pre-divider text never contains a | character, this is the cleanest, most performant option. We directly match all characters that aren't | until we hit the separator:
grammar HC { token TOP { <pre> <divider> <post> } token pre { <-[|]>* } # Match any character except |, zero or more times token divider { '|' ** 1..2 } # 1 or 2 pipe characters token post { .* } # Match everything after the divider }
This works perfectly for inputs like Want this || Not this or Want this | Not this, capturing the full pre-separator text as a single <pre> node.
Solution 2: Lookahead for Flexible Pre-Text (If Pre-Text Can Include |)
If your pre-divider text might have single | characters (e.g., Hello | world || foo), use a lookahead to ensure we only stop matching <pre> when we're about to hit the 1-2 pipe separator:
grammar HC { token TOP { <pre> <divider> <post> } token pre { .*? <?before <divider>> } # Match until we see the divider coming up token divider { '|' ** 1..2 } token post { .* } }
The <?before <divider>> assertion tells the parser: "stop matching <pre> right before you encounter the divider pattern". This lets you have single pipes in the pre-text while still correctly identifying the actual separator.
Key Notes
- Raku's
tokenis inherently:ratchet(no backtracking), so these grammars are efficient and won't waste cycles re-parsing. - Use
\Ninstead of.in<pre>or<post>if you want to exclude newline characters (match only single-line text).
内容的提问来源于stack exchange,提问作者StevieD

