ANTLR入门遇变量声明规则匹配错误求助:mismatched input 'var'问题排查
Hey there! Let's break down why your var test; input is throwing that confusing mismatched input error, and fix it step by step.
The Root Cause
That error message looks contradictory at first ("expecting 'var' but got 'var'")—but the problem is happening before the parser even gets to work. ANTLR's tokenizer is misclassifying your var keyword as a STRING instead of the VAR token your parser expects.
Here's why:
- Your
STRINGrule is defined as[a-zA-Z]+—this matches any sequence of letters, including keywords likevarandif. - In ANTLR 4, when multiple token rules can match the same input, the first defined rule takes priority. Since
STRINGcomes beforeVARandIFin your grammar, those keywords get swallowed up asSTRINGtokens instead of their dedicated keyword tokens. - On top of that, your
NAMErule is also[a-zA-Z]+, which creates another conflict with bothSTRINGand your keywords.
Step-by-Step Fixes
1. Fix the STRING Token Definition
First, strings should be enclosed in quotes (this is standard for most languages). Update your STRING rule to match quoted text instead of plain letters:
STRING: '"' [a-zA-Z0-9_]* '"' | '\'' [a-zA-Z0-9_]* '\'';
This will match strings like "test" or 'hello_world' instead of conflicting with identifiers and keywords.
2. Reorder Token Rules
Move all keyword tokens (like VAR, IF, BOOLEAN) before your NAME rule. ANTLR will prioritize these exact matches over the general NAME rule:
/* Tokens - reordered to prioritize keywords first */ VAR : 'var'; IF: 'if'; BOOLEAN: 'true' | 'false'; NAME : [a-zA-Z_][a-zA-Z0-9_]*; // Updated to allow underscores and start with letter/underscore (standard for identifiers) NUM : '0' | '-'?[1-9][0-9]*; STRING: '"' [a-zA-Z0-9_]* '"' | '\'' [a-zA-Z0-9_]* '\''; SEMICOLON : ';'; LPAREN: '('; RPAREN: ')'; LCURLY: '{'; RCURLY: '}'; EQUALITY: '==' | '<' | '>' | '<=' | '>=' | '!=' ; operation: '+' | '-' | '*' | '/'; operand: NUM | NAME; // Added NAME as a valid operand (makes sense for expressions) WS : [ \t\r\n]+ -> skip;
3. Clean Up Ambiguous Rules
Your term rule has an empty alternative (term: | ...), which can lead to parsing ambiguity. If you want to allow optional terms, use a ? modifier instead:
term: (NUM EQUALITY NUM | NAME EQUALITY NUM | NUM EQUALITY NAME | NAME EQUALITY NAME)?;
Or remove the empty option entirely if terms should always have a comparison.
Full Fixed Grammar
Here's the complete corrected grammar:
grammar treetwo; program : (declaration | statement)+ EOF; declaration : variable_declaration | variable_assignment ; statement: expression | ifstmnt ; variable_declaration: VAR NAME SEMICOLON ; variable_assignment: NAME '=' NUM SEMICOLON | NAME '=' STRING SEMICOLON | NAME '=' BOOLEAN SEMICOLON ; expression: operand operation operand SEMICOLON | expression operation expression SEMICOLON | operand operation expression SEMICOLON | expression operation operand SEMICOLON ; ifstmnt: IF LPAREN term RPAREN LCURLY (declaration | statement)+ RCURLY ; term: (NUM EQUALITY NUM | NAME EQUALITY NUM | NUM EQUALITY NAME | NAME EQUALITY NAME)?; /* Tokens - reordered and fixed */ VAR : 'var'; IF: 'if'; BOOLEAN: 'true' | 'false'; NAME : [a-zA-Z_][a-zA-Z0-9_]*; NUM : '0' | '-'?[1-9][0-9]*; STRING: '"' [a-zA-Z0-9_]* '"' | '\'' [a-zA-Z0-9_]* '\''; SEMICOLON : ';'; LPAREN: '('; RPAREN: ')'; LCURLY: '{'; RCURLY: '}'; EQUALITY: '==' | '<' | '>' | '<=' | '>=' | '!=' ; operation: '+' | '-' | '*' | '/'; operand: NUM | NAME; WS : [ \t\r\n]+ -> skip;
Testing the Fix
Now when you input var test;, the tokenizer will correctly identify var as the VAR token, and test as a NAME token. The parser will match this to your variable_declaration rule without errors.
内容的提问来源于stack exchange,提问作者Gerard Colman

