JavaCC自定义错误引发空串匹配问题及自定义错误消息实现咨询
I’ve dealt with this exact frustration in JavaCC before—trying to add specific error messages but hitting that annoying empty string match error. Let’s break down why this happens and how to fix it properly.
Why the Error Occurs
Your initial attempt to add a throw statement as a fallback branch in A() or B() creates an empty expansion (a branch that doesn’t match or consume any tokens). Since your Foo() rule uses (A() B())+ (a one-or-more repetition), JavaCC flags this because it doesn’t allow repeatable units to be matchable by an empty string—this would lead to infinite loop risks and ambiguous parsing.
Correct Implementation
The key is to ensure every branch in your non-terminal rules either matches a valid token or consumes an invalid one (to avoid loops), so no empty expansions exist. Here’s the revised syntax:
Full Working Grammar
Foo(): {} { (A() B())+ } A(): {} { <TOKA1> | <TOKA2> | { // Capture the invalid token Token badToken = getToken(1); // Consume the bad token to prevent infinite parsing loops consume(); // Throw a custom, specific error throw new ParseException( String.format("Expected an A-type token (TOKA1/TOKA2), but found '%s' at line %d, column %d", badToken.image, badToken.beginLine, badToken.beginColumn) ); } } B(): {} { <TOKB1> | <TOKB2> | { Token badToken = getToken(1); consume(); throw new ParseException( String.format("Expected a B-type token (TOKB1/TOKB2), but found '%s' at line %d, column %d", badToken.image, badToken.beginLine, badToken.beginColumn) ); } }
Key Fixes Explained
- No empty branches: The error-handling branch now calls
consume()to eat the invalid token, so this branch isn’t an empty expansion anymore. JavaCC won’t throw the "empty string match" error. - Specific error details: We capture the invalid token’s image and position to create a far more useful error message than JavaCC’s generic default.
- Avoids infinite loops: Without
consume(), the parser would keep trying to matchA()/B()against the same invalid token forever.
Additional Tips
- If you want to handle errors without consuming the token (e.g., for multi-token error recovery), use a
LOOKAHEADcheck to validate that the next token isn’t a valid match before entering the error branch:A(): {} { <TOKA1> | <TOKA2> | LOOKAHEAD(!(<TOKA1> | <TOKA2>)) { // Same error handling as above } } - For more structured error handling, create a custom subclass of
ParseExceptionthat includes error codes or additional context.
内容的提问来源于stack exchange,提问作者Daniel Causebrook

