You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Antlr4语法修复:适配含空格值列表的内部语言解析需求

Fixing ANTLR4 Grammar for List Items with Spaces

Got it, the problem here is your current ANTLR4 grammar can't handle list items or identifiers that include spaces—like Mechanical Instability in your $IF clause or Type grammar Trigger after $FROM. Let's walk through why this breaks and how to fix it:

Why Your Current Grammar Fails

Your existing rules have two key limitations:

  • The IDENTIFIER rule only allows [a-zA-Z_.'0-9]+, which excludes spaces. This means phrases with spaces get split into multiple IDENTIFIER tokens instead of being treated as a single value.
  • LIST_VALUES is built on IDENTIFIER, so it can't recognize list items with spaces either.

Modified ANTLR4 Grammar

Here's the updated grammar that fixes these issues, while preserving all your original language features:

prog: expr EOF;

expr : IF (ANY|EMPTY|NOT|comparator|count|complex|count_complex|LIST_ITEM|LIST_VALUES) FROM IDENTIFIER_WITH_SPACES THEN 
     | expr binary expr 
     | LPAREN expr RPAREN 
     ;

comparator : (GT | GE | LT | LE | EQ) (INT|IDENTIFIER_WITH_SPACES);
binary : AND | OR;
count : COUNT LPAREN (IDENTIFIER_WITH_SPACES | DISTINCT) RPAREN comparator;
complex : COMPLEX LPAREN expr RPAREN;
count_complex : COUNT_COMPLEX LPAREN (expr | DISTINCT IDENTIFIER_WITH_SPACES expr) RPAREN comparator;

// Keyword rules (higher priority than text tokens)
IF : '$IF';
FROM : '$FROM';
THEN : '$THEN';
AND : '$AND' ;
ANY : '$ANY';
EMPTY : '$EMPTY';
DISTINCT : '$DISTINCT';
COUNT : '$COUNT';
COUNT_COMPLEX : '$COUNT_COMPLEX';
COMPLEX : '$COMPLEX';
OR : '$OR' ;
NOT : '$NOT' LPAREN LIST_ITEM (',' LIST_ITEM)*? RPAREN;

// Comparison operators
GT : '>' ;
GE : '>=' ;
LT : '<' ;
LE : '<=' ;
EQ : '=' ;

// Punctuation
LPAREN : '(' ;
RPAREN : ')' ;

// Numeric values
INT : '-'?[0-9]+;

// Identifiers that can include spaces (for $FROM targets)
IDENTIFIER_WITH_SPACES : ~[$ \r\t\u000C\n]+ ( [ ] ~[$ \r\t\u000C\n]+ )*;

// List items that support spaces (stops at commas or $-keywords)
LIST_ITEM : ~[,$ \r\t\u000C\n]+ ( [ ] ~[,$ \r\t\u000C\n]+ )*;

// Comma-separated list values
LIST_VALUES : LIST_ITEM ( ',' LIST_ITEM )*;

// Skip whitespace
WS : [ \r\t\u000C\n]+ -> skip;

Key Changes Explained

  • IDENTIFIER_WITH_SPACES: Replaces your original IDENTIFIER for $FROM targets. It matches text with spaces, but stops at $ to avoid swallowing keywords like $THEN.
  • LIST_ITEM: Designed to match list entries with spaces, stopping at commas or $-prefixed keywords. This lets Mechanical Instability be treated as a single list item.
  • Updated LIST_VALUES: Now uses LIST_ITEM instead of IDENTIFIER, so it correctly parses comma-separated entries with spaces.
  • Adjusted NOT Rule: Swapped CHAR for LIST_ITEM to let $NOT handle values with spaces too (if your language requires it).
  • Keyword Priority: All $-prefixed keywords are defined before text tokens, ensuring ANTLR4 prioritizes them over longer text matches (so $FROM isn't mistaken for part of an identifier).

Test with Your Example

Your sample statement $IF Mechanical Instability,Deformity $FROM Type grammar Trigger $THEN will now parse correctly:

  • Mechanical Instability and Deformity are recognized as two distinct LIST_ITEM tokens in a LIST_VALUES
  • Type grammar Trigger is parsed as a single IDENTIFIER_WITH_SPACES

内容的提问来源于stack exchange,提问作者IuryRibeiro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:56:03