Antlr4语法修复:适配含空格值列表的内部语言解析需求
Fixing ANTLR4 Grammar for List Items with Spaces
Got it, the problem here is your current ANTLR4 grammar can't handle list items or identifiers that include spaces—like Mechanical Instability in your $IF clause or Type grammar Trigger after $FROM. Let's walk through why this breaks and how to fix it:
Why Your Current Grammar Fails
Your existing rules have two key limitations:
- The
IDENTIFIERrule only allows[a-zA-Z_.'0-9]+, which excludes spaces. This means phrases with spaces get split into multipleIDENTIFIERtokens instead of being treated as a single value. LIST_VALUESis built onIDENTIFIER, so it can't recognize list items with spaces either.
Modified ANTLR4 Grammar
Here's the updated grammar that fixes these issues, while preserving all your original language features:
prog: expr EOF; expr : IF (ANY|EMPTY|NOT|comparator|count|complex|count_complex|LIST_ITEM|LIST_VALUES) FROM IDENTIFIER_WITH_SPACES THEN | expr binary expr | LPAREN expr RPAREN ; comparator : (GT | GE | LT | LE | EQ) (INT|IDENTIFIER_WITH_SPACES); binary : AND | OR; count : COUNT LPAREN (IDENTIFIER_WITH_SPACES | DISTINCT) RPAREN comparator; complex : COMPLEX LPAREN expr RPAREN; count_complex : COUNT_COMPLEX LPAREN (expr | DISTINCT IDENTIFIER_WITH_SPACES expr) RPAREN comparator; // Keyword rules (higher priority than text tokens) IF : '$IF'; FROM : '$FROM'; THEN : '$THEN'; AND : '$AND' ; ANY : '$ANY'; EMPTY : '$EMPTY'; DISTINCT : '$DISTINCT'; COUNT : '$COUNT'; COUNT_COMPLEX : '$COUNT_COMPLEX'; COMPLEX : '$COMPLEX'; OR : '$OR' ; NOT : '$NOT' LPAREN LIST_ITEM (',' LIST_ITEM)*? RPAREN; // Comparison operators GT : '>' ; GE : '>=' ; LT : '<' ; LE : '<=' ; EQ : '=' ; // Punctuation LPAREN : '(' ; RPAREN : ')' ; // Numeric values INT : '-'?[0-9]+; // Identifiers that can include spaces (for $FROM targets) IDENTIFIER_WITH_SPACES : ~[$ \r\t\u000C\n]+ ( [ ] ~[$ \r\t\u000C\n]+ )*; // List items that support spaces (stops at commas or $-keywords) LIST_ITEM : ~[,$ \r\t\u000C\n]+ ( [ ] ~[,$ \r\t\u000C\n]+ )*; // Comma-separated list values LIST_VALUES : LIST_ITEM ( ',' LIST_ITEM )*; // Skip whitespace WS : [ \r\t\u000C\n]+ -> skip;
Key Changes Explained
IDENTIFIER_WITH_SPACES: Replaces your originalIDENTIFIERfor$FROMtargets. It matches text with spaces, but stops at$to avoid swallowing keywords like$THEN.LIST_ITEM: Designed to match list entries with spaces, stopping at commas or$-prefixed keywords. This letsMechanical Instabilitybe treated as a single list item.- Updated
LIST_VALUES: Now usesLIST_ITEMinstead ofIDENTIFIER, so it correctly parses comma-separated entries with spaces. - Adjusted
NOTRule: SwappedCHARforLIST_ITEMto let$NOThandle values with spaces too (if your language requires it). - Keyword Priority: All
$-prefixed keywords are defined before text tokens, ensuring ANTLR4 prioritizes them over longer text matches (so$FROMisn't mistaken for part of an identifier).
Test with Your Example
Your sample statement $IF Mechanical Instability,Deformity $FROM Type grammar Trigger $THEN will now parse correctly:
Mechanical InstabilityandDeformityare recognized as two distinctLIST_ITEMtokens in aLIST_VALUESType grammar Triggeris parsed as a singleIDENTIFIER_WITH_SPACES
内容的提问来源于stack exchange,提问作者IuryRibeiro
相关产品推荐
相关产品推荐

