You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更简便地自定义ANTLR4各类错误提示消息?

嘿,你的现有实现已经很合理了!不过确实有更简洁的方式来自定义这些语法错误消息,不用分别重写Lexer的notifyListeners和DefaultErrorStrategy的一堆report方法,下面给你两种更省心的方案:

方案1:用自定义ANTLRErrorListener统一处理所有错误

这是最简洁的方式——直接实现自己的错误监听器,在里面根据异常类型生成对应的自定义消息,然后替换掉默认的监听器即可。这样所有错误逻辑都集中在一个类里,不用去修改Lexer和Parser的错误策略类。

比如这个自定义监听器的示例:

import org.antlr.v4.runtime.*;
import org.antlr.v4.runtime.misc.Interval;

public class CustomErrorListener extends BaseErrorListener {
    @Override
    public void syntaxError(Recognizer<?, ?> recognizer, Object offendingSymbol, int line, int charPositionInLine, String defaultMsg, RecognitionException e) {
        String customMsg;

        // 针对词法错误(LexerNoViableAltException)
        if (e instanceof LexerNoViableAltException) {
            Lexer lexer = (Lexer) recognizer;
            String invalidText = lexer._input.getText(Interval.of(lexer._tokenStartCharIndex, lexer._input.index()));
            customMsg = String.format("Operator '%s' is unknown.", lexer.getErrorDisplay(invalidText));
        }
        // 针对NoViableAlternative错误
        else if (e instanceof NoViableAltException) {
            NoViableAltException nvae = (NoViableAltException) e;
            TokenStream tokens = ((Parser) recognizer).getInputStream();
            String inputSegment;

            if (tokens != null) {
                if (nvae.getStartToken().getType() == Token.EOF) {
                    inputSegment = "<EOF>";
                } else {
                    inputSegment = tokens.getText(nvae.getStartToken(), nvae.getOffendingToken());
                }
            } else {
                inputSegment = "<unknown input>";
            }
            customMsg = String.format("Invalid operation: '%s'.", inputSegment);
        }
        // 针对InputMismatch错误
        else if (e instanceof InputMismatchException) {
            Token offendingToken = (Token) offendingSymbol;
            customMsg = String.format("Unexpected token '%s' at position %d:%d. Expected a different type here.",
                    offendingToken.getText(), line, charPositionInLine);
        }
        // 针对UnwantedToken错误
        else if (e instanceof UnwantedTokenException) {
            UnwantedTokenException ute = (UnwantedTokenException) e;
            customMsg = String.format("Unexpected token '%s' - please remove it.", ute.getOffendingToken().getText());
        }
        // 针对MissingToken错误
        else if (e instanceof MissingTokenException) {
            MissingTokenException mte = (MissingTokenException) e;
            String expectedTokenType = ((Parser) recognizer).getVocabulary().getDisplayName(mte.getExpectedTokenType());
            customMsg = String.format("Missing required token '%s' at position %d:%d.", expectedTokenType, line, charPositionInLine);
        }
        // 兜底用默认消息
        else {
            customMsg = defaultMsg;
        }

        // 输出错误(你也可以改成自己的日志/处理逻辑)
        System.err.printf("Syntax Error [Line %d:%d]: %s%n", line, charPositionInLine, customMsg);
    }
}

然后只需要在初始化Lexer和Parser的时候替换默认监听器:

// 替换Lexer的错误监听器
MyCustomLexer lexer = new MyCustomLexer(inputStream);
lexer.removeErrorListeners(); // 移除默认的ConsoleErrorListener
lexer.addErrorListener(new CustomErrorListener());

// 替换Parser的错误监听器
CommonTokenStream tokens = new CommonTokenStream(lexer);
MyCustomParser parser = new MyCustomParser(tokens);
parser.removeErrorListeners();
parser.addErrorListener(new CustomErrorListener());
方案2:复用DefaultErrorStrategy仅修改消息

如果你想保留DefaultErrorStrategy自带的错误恢复逻辑(比如跳过错误继续解析),只是想替换错误消息,那么可以只重写对应的report方法,但简化代码——比如复用父类里获取输入片段的逻辑,不用自己再写一遍:

import org.antlr.v4.runtime.*;
import org.antlr.v4.runtime.atn.NoViableAltException;

public class CustomErrorStrategy extends DefaultErrorStrategy {
    @Override
    protected void reportNoViableAlternative(Parser recognizer, NoViableAltException e) {
        String input = getTokenErrorDisplay(e.getOffendingToken());
        // 如果是EOF的情况单独处理
        if (e.getStartToken().getType() == Token.EOF) {
            input = "<EOF>";
        } else {
            TokenStream tokens = recognizer.getInputStream();
            if (tokens != null) {
                input = tokens.getText(e.getStartToken(), e.getOffendingToken());
            }
        }
        String msg = String.format("Invalid operation: '%s'.", input);
        recognizer.notifyErrorListeners(e.getOffendingToken(), msg, e);
    }

    @Override
    protected void reportInputMismatch(Parser recognizer, InputMismatchException e) {
        Token offendingToken = e.getOffendingToken();
        String msg = String.format("Unexpected token '%s' - expected one of: %s.",
                getTokenErrorDisplay(offendingToken),
                e.getExpectedTokens().toString(recognizer.getVocabulary()));
        recognizer.notifyErrorListeners(offendingToken, msg, e);
    }

    // 同理重写reportUnwantedToken、reportMissingToken等方法
}

然后给Parser设置这个自定义策略:

parser.setErrorHandler(new CustomErrorStrategy());
哪种方案更适合?
  • 如果你想统一管理所有错误消息和处理逻辑,方案1更优——所有错误都在一个类里处理,不用分散到Lexer和ErrorStrategy中。
  • 如果你依赖DefaultErrorStrategy的错误恢复行为(比如不想因为一个错误就终止解析),方案2更合适,只修改消息保留原有恢复逻辑。

内容的提问来源于stack exchange,提问作者mrx8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:03:48