You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在ANTLR自定义ErrorListener中获取错误符号信息?

解决ANTLR解析未知符号时错误信息定位问题

在ANTLR中,词法错误(比如遇到未知符号#)和语法错误的处理逻辑存在差异。默认的BaseErrorListener处理词法错误时,offendingSymbol不会指向错误符号本身,而是后续的有效符号——这是因为Lexer遇到无法匹配规则的字符时,会先消耗该字符再触发错误,导致输入流指针已经移动到下一个位置。

要正确获取引发词法错误的符号,需在自定义ErrorListener中区分Lexer和Parser场景:

  • 当recognizer是Lexer实例时,判定为词法错误,直接从输入流回退提取错误字符
  • 当recognizer是Parser实例时,按照常规语法错误逻辑处理即可

以下是完整的自定义ErrorListener实现代码:

import org.antlr.v4.runtime.*;

public class CustomErrorListener extends BaseErrorListener {
    @Override
    public void syntaxError(Recognizer<?, ?> recognizer, Object offendingSymbol, int line, int charPositionInLine, String msg, RecognitionException e) {
        String errorMessage;
        if (recognizer instanceof Lexer) {
            // 处理词法错误:提取未知符号
            CharStream charStream = recognizer.getInputStream();
            int errorCharIndex = charStream.index() - 1;
            String errorChar = charStream.getText(Interval.of(errorCharIndex, errorCharIndex));
            errorMessage = String.format("line %d:%d token recognition error at: '%s'", line, charPositionInLine, errorChar);
        } else {
            // 处理语法错误:使用 offendingSymbol 信息
            Token token = (Token) offendingSymbol;
            errorMessage = String.format("line %d:%d syntax error at token '%s'", line, charPositionInLine, token.getText());
        }
        System.err.println(errorMessage);
    }
}

测试公式2 # 4时,控制台会输出预期的错误信息:

line 1:2 token recognition error at: '#'

内容的提问来源于stack exchange,提问作者zahaand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 14:00:05