如何修改ANTLR嵌套查询解析规则以支持多匹配括号?
Alright, let's figure out how to fix your ANTLR rule so it handles balanced parentheses properly. The core issue with your current rule is that it doesn’t track matching opening/closing bracket pairs—when it hits the first ), it stops parsing instead of recognizing the full balanced group. Here’s a revised approach that fixes this:
Revised ANTLR Rules
nested_query: non_paren_content | nested_query nested_query_op LPARAN nested_query RPARAN | LPARAN nested_query RPARAN | nested_query non_paren_content ; non_paren_content: (~(LPARAN | RPARAN))+; nested_query_op: binary_in | binary_not_in ; binary_in: 'in'; binary_not_in: 'not' 'in'; LPARAN: '('; RPARAN: ')';
Key Changes & Explanations
Dedicated
non_paren_contentrule
We replaced the inline~(LPARAN | RPARAN)+?with a separate rule that matches any sequence of non-bracket characters. This avoids the non-greedy match issue that caused early termination when encountering a closing bracket.Explicit balanced bracket branch
AddingLPARAN nested_query RPARANas a top-level branch lets the parser recursively handle nested bracket groups. This means it will correctly parse content like(typeTag ="TAG_SRC_IP_VM")or even deeper nesting like((a in (b in (c)))).Flexible sequence handling
Thenested_query non_paren_contentbranch ensures that text can appear before or after nested bracket groups (like your examplelist(srcVm) of flows...wherelistis non-bracket content followed by a bracket-enclosed group).
Testing with Your Example
For the input list(srcVm) of flows where (typeTag ="TAG_SRC_IP_VM") until timestamp, the parser will now:
- Recognize
list(srcVm)asnon_paren_content+LPARAN nested_query RPARAN(wheresrcVmisnon_paren_content) - Correctly parse
(typeTag ="TAG_SRC_IP_VM")as a balanced bracket group containing non-bracket content - Handle the rest of the text seamlessly as
non_paren_content
This rule will work for any number of balanced nested parentheses while preserving your original in/not in operator logic.
内容的提问来源于stack exchange,提问作者tuk

